Fluidstack · Operations · Unspecified · Posted 2026-07-17
Production Engineer, Facilities
Fluidstack · San Francisco, CA; New York, NY; Austin, TX; Seattle, WA · $173k–279k base
This range's midpoint is above 72% of posted operations ranges at AI companies right now. See the salary index.
Apply on Fluidstack's site Watch Fluidstack for new roles
ABOUT FLUIDSTACK
We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.
We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.
We hire people who care deeply about this problem space. If that is you, please apply!
HOW WE OPERATE
- Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.
- Insane urgency. We drive everything forward as fast as possible.
- Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.
- Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.
- Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.
THE PRODUCTION ENGINEERING TEAM
Examples of key problems the team is working on
- Scaling the systems which make the physical plant observable and operable as one fleet. Building the telemetry, alarm, topology, and health systems that allow operators and automations to see the true state of power, cooling, and environmental infrastructure across every site.
- Turn facility incidents into a closed repair loop. Automate the path from detection and diagnosis through maintenance, remediation, validation, and return to service so failures do not disappear into handoffs between software, engineering, vendors, and site operations.
- Bring new sites and equipment into production safely at construction speed. Build repeatable readiness gates, commissioning signals, staged deployments, canaries, and rollback mechanisms for a fleet growing by multiple sites at once.
- Keep operators ahead of power and cooling risk. Build capacity views, safeguards, anomaly detection, service-health reviews, and operational tooling that identify problems before they affect customers.
ROLE SCOPE
- Carry the facilities production on-call pager and lead incidents involving facility software, telemetry, controls integrations, and automation. Diagnose the failure, coordinate the responsible teams, restore service, and drive the systemic fix.
- Own the production reliability of the facilities telemetry and alarm platform end to end. Build and operate ingestion, storage, APIs, data-quality checks, actionable alerts, retention, backups, failover, and recovery across industrial protocols and site integrations.
- Turn diagnosis and repair into pipelines rather than procedures. Build Python or Go tooling for fleet-wide debugging, maintenance workflows, automated validation, incident response, and safe return to service.
- Own production deployment and runtime management for facilities services, including BMS and EPMS integrations, SCADA platforms such as Ignition, virtual PLCs, demand management.
- Define and enforce production-readiness standards for new sites, equipment, APIs, telemetry integrations, and controls deployments. Build the tests, canaries, release gates, staged promotion, and rollback mechanisms that define what healthy looks like before launch.
- Own the operational maturity of every in-scope service. Establish SLOs, capacity plans, health dashboards, runbooks …
More operations roles at Fluidstack
-
Construction Manager
OperationsLead / ManagerRemote US$155k–210ktoday
-
Construction Project Manager
OperationsRemote US$200k–253ktoday
-
Pre-Construction Manager
OperationsLead / ManagerUS$186k–240ktoday
-
Facilities Manager
OperationsLead / ManagerUS$110k–164k3d
-
Environmental Manager, HSE
OperationsLead / ManagerUS$150k–185k11d
See also: Systems Engineer jobs · AI jobs in San Francisco Bay Area · Fluidstack salaries · Python jobs · Kubernetes jobs.
This listing is reproduced from Fluidstack's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All Fluidstack roles · AI salaries.