AI Hiring Index

Snorkel AI · Engineering · Staff+ · Posted 2026-08-27

Senior/Staff FDE - Synthetic Data Generation

Snorkel AI · New York City, NY (Hybrid); San Francisco, CA (Hybrid) · $180k–320k base

This range's midpoint is above 65% of posted engineering ranges at AI companies right now. See the salary index.

Apply on Snorkel AI's site Watch Snorkel AI for new roles

About Snorkel

Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine technology with research-driven AI data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes. 

Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!

About the Role

Snorkel AI is hiring a Forward Deployed Engineer focused on Synthetic Data Generation to partner with leading AI labs and enterprises on their most critical AI initiatives.

In this role, you will lead the technical execution of complex customer engagements where synthetic data is used to improve model training, evaluation, and performance. You will translate ambiguous model and data challenges into effective data strategies, build scalable generation and evaluation pipelines, and use experimentation to continuously improve data quality and downstream model outcomes.

You will work across the full delivery lifecycle—from technical discovery and solution design through implementation, evaluation, and production delivery. You will also identify patterns across engagements and turn successful approaches into reusable capabilities, technical standards, and product improvements.

Main Responsibilities

Synthetic Data Generation & Evaluation

Design and build scalable synthetic data generation, transformation, filtering, and evaluation pipelines for complex AI use cases

Translate model objectives, failure modes, and data gaps into synthetic data strategies, experiments, and technical specifications

Develop LLM- and ML-assisted workflows to generate high-quality training and evaluation datasets across targeted behaviors, domains, and edge cases

Build automated evaluators, quality checks, and measurement frameworks to assess correctness, relevance, diversity, coverage, and adherence to customer requirements

Design and run experiments to measure the impact of synthetic data on downstream model performance and iteratively improve generation approaches

Package and deliver production-grade datasets with standardized formats, quality assurance, and clear documentation

Forward Deployed Engineering & Customer Partnership

Lead technical workstreams from initial solution design through production delivery, navigating ambiguity and making sound technical decisions

Build, refine, and iterate on solutions that address customer needs, incorporating feedback to ensure the delivered work provides tangible value

Rapidly prototype and productionize solutions across models, data pipelines, APIs, and custom applications

Communicate technical tradeoffs, experimental results, and recommendations clearly to technical and cross-functional stakeholders

Serve as a trusted technical partner to customers and internal delivery teams, resolving complex blockers and driving alignment

Technical Leadership & Scale

Identify recurring patterns across customer engagements and turn successful solutions into reusable pipelines, evaluators, tooling, and best practices

Define and improve technical standards for synthetic data generation, experimentation, evaluation, and delivery

Partner with DaaS Engineering and Product teams to influence platform and product capabilities based on real-world customer needs

Lead technical design reviews, share expertise, and provide guidance to other engineers

Stay current with emerging synthetic data, LLM evaluation, and data curation techniques and assess their applicability to customer problems

What We're Looking For

5+ years of experience in machine learning engineering, data science, applied AI, forward deployed engineering, or a similar technical role

Strong Python skills and experience building reliable production data or ML systems, including containerizing …

New engineering roles at AI companies, every Monday. The week's openings in this function across 286 companies, plus the weekly index. Free.

More engineering roles at Snorkel AI

See also: AI jobs in San Francisco Bay Area · Snorkel AI salaries · Python jobs · Docker jobs · AWS jobs.

This listing is reproduced from Snorkel AI's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All Snorkel AI roles · AI salaries.