OpenAI · Infrastructure · Unspecified · Posted 2026-09-28
Software Engineer, DevOps
OpenAI · San Francisco · $177k–327k base
This range's midpoint is above 72% of posted infrastructure ranges at AI companies right now. See the salary index.
Apply on OpenAI's site Watch OpenAI for new roles
About the Team
OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform.
About the Role
You will define and build the CI/CD systems that qualify frontier models and the software stack running on OpenAI’s custom AI accelerators. Your work will make correctness, performance, and stability measurable on every change, from individual commits through production releases.
This role spans model workloads, accelerator software, systems infrastructure, and developer productivity. You will establish scalable CI/CD, GitOps, and monorepo practices; design trustworthy regression and benchmarking pipelines; and partner with model, compiler, kernel, runtime, firmware, and hardware teams to shorten feedback loops without compromising production readiness.
In this role, you will:
- Define the model CI strategy for frontier models running on OpenAI custom silicon and integrated accelerator systems.
- Build production CI/CD pipelines for functional regression, performance benchmarking, software stability, and release qualification.
- Design representative test matrices across models, configurations, hardware generations, software components, and deployment environments.
- Create reliable performance baselines, regression detection, bisect and triage workflows, and clear ownership for failures.
- Establish GitOps practices for reproducible configuration, promotion, rollback, auditability, and environment consistency.
- Shape monorepo architecture, dependency management, build and test boundaries, change validation, and developer workflows at scale.
- Develop scalable orchestration, artifact management, caching, scheduling, observability, and capacity controls for accelerator-backed CI.
- Partner with model, compiler, kernel, runtime, firmware, validation, and hardware teams to translate release risks into automated gates.
- Improve CI reliability, speed, debuggability, and cost efficiency while maintaining high confidence in production software.
You might thrive in this role if:
- Have built or operated large-scale CI/CD, developer infrastructure, test automation, or production engineering systems.
- Understand modern software delivery practices, including GitOps, release promotion, rollback, reproducibility, and policy-driven automation.
- Have experience designing monorepos, build systems, dependency graphs, test selection, or change-impact analysis.
- Can design regression and performance-benchmarking systems with stable baselines, useful metrics, and actionable failure diagnosis.
- Are proficient in Python, Go, Rust, C++, or another language used to build reliable infrastructure and automation.
- Have worked with distributed systems, schedulers, containers, clusters, or heterogeneous compute infrastructure.
- Can collaborate with model and systems engineers to turn complex accelerator workloads into repeatable production qualification.
- Care deeply about reliability, observability, developer experience, and reducing time from code change to trustworthy signal.
To comply with U.S. export control laws and regulations, candidates for this role may need to meet certain legal status requirements as provided in those laws and regulations.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human n …
More infrastructure roles at OpenAI
-
Platform Engineering Manager, Forward Deployed Engineering (FDE)
InfrastructureLead / ManagerRemote US$302k–335ktoday
-
Data Center Hardware Quality & Reliability Engineer
InfrastructureStaff+Remote US$226k–285ktoday
-
Technical Program Manager, Hardware Systems
InfrastructureUS$207k–242ktoday
-
Software Engineer, Search Infrastructure
InfrastructureRemote US$266k–445k3d
-
Rack Power Engineer
InfrastructureStaff+Remote US$287k–485k5d
See also: Infrastructure Engineer jobs · AI jobs in San Francisco Bay Area · OpenAI salaries · Python jobs · C++ jobs · Rust jobs.
This listing is reproduced from OpenAI's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All OpenAI roles · AI salaries.