Staff Software Engineer, Search & Retrieval Infrastructure
New
S
SylloLegal Technology
United States - RemoteFull-TimeStaff
Salary$190,000 — $230,000 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years
- Required Skills
- PythonElasticSearchGCPKafkaKubernetesGoRustDistributed Systems
Requirements
- 8+ years of software engineering experience, with a proven track record operating at the Staff/Principal level.
- Deep, production-level expertise tuning and scaling Lucene-based search engines (Elasticsearch, Solr) and modern vector indexing infrastructure.
- Proven experience optimizing and scaling highly distributed, high-throughput systems to handle petabyte-level data.
- Deep understanding of index internals, chunking strategies, and embedding retrieval optimization.
- Strong history of managing compute vs. storage trade-offs and designing cost-effective architectures.
- Extensive experience managing complex data pipelines and high-throughput event streaming (Kafka, Kinesis).
- Expert command of cloud primitives (GCP preferred), Kubernetes, and infrastructure-as-code.
- Expert-level proficiency in systems-level and backend languages (Go, Rust, Python, or Java/C++).
Responsibilities
- Scale the Retrieval Stack: Lead the optimization and architectural evolution of our existing hybrid search infrastructure, maximizing the throughput and efficiency of both lexical search (e.g., Elasticsearch, Lucene) and dense vector databases.
- Advanced Data Tiering & Scanning: Design and implement intelligent, cost-effective tiering strategies across hot, warm, and cold data states.
- Evolve our distributed pipelines to efficiently execute asynchronous, massive-scale scans of petabytes of data in varying states of availability.
- Relentless Optimization: Drive down latency and cost-to-serve by analyzing system bottlenecks and tuning indexing and querying algorithms.
- Technical Leadership: Act as the domain expert and owner of the indexing and search ecosystem, setting the long-term technical vision.
- Resiliency at Scale: Ensure fault-tolerant, highly available operations during massive parallel ingest events and complex, concurrent querying.
View Full Description & ApplyYou'll be redirected to the employer's site