Vast.ai · Infrastructure · Unspecified · Posted 2026-07-30
AI, HPC & GPU Infrastructure Support Engineer
Vast.ai · Los Angeles · $90k–150k base
This range sits in the bottom 5% of posted infrastructure ranges at AI companies right now. See the salary index.
Apply on Vast.ai's site Watch Vast.ai for new roles
About Us
Vast.ai http://Vast.ai's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation.
We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team.
About the Role
This role focuses on troubleshooting complex Linux and GPU infrastructure issues across NVIDIA drivers, CUDA, GPU workloads, Ubuntu, Docker, KVM based virtual machines, networking, hardware, BIOS, and firmware. You’ll investigate failures, reproduce issues, identify root causes, and propose practical solutions across the full infrastructure stack.
You’ll also serve as the engineering resource our L1 support team relies on when tickets go beyond frontline triage. You’ll own complex escalations end-to-end, gather technical evidence, coordinate with the appropriate teams, and communicate findings clearly to clients, infrastructure suppliers, and internal teams.
The best engineers in this role don’t just resolve individual issues—they recognize recurring patterns, improve diagnostic tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic Linux, GPU, and infrastructure problems.
Strong GPU troubleshooting experience, Linux systems knowledge, and technical support skills are the primary requirements. You should be comfortable working autonomously in Ubuntu environments and troubleshooting NVIDIA drivers, CUDA, containers, virtual machines, networking, hardware, and GPU workloads.
Vast.ai http://Vast.ai users or hosts strongly preferred.
LOCATION AND SCHEDULE
This is a full-time position based in our Westwood, Los Angeles office.
Available schedules:
- Monday–Friday: Fully on-site
- Sunday–Thursday: Four days on-site and one day working from home
Key Responsibilities
- Diagnose and resolve issues across NVIDIA CUDA/GPU drivers, Docker, and KVM virtualization environments
- Investigate GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O bottlenecks
- Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU-accelerated workloads
- Troubleshoot network-layer issues, including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines
- Handle escalated support tickets involving GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration
- Provide managed support for supplier onboarding and ongoing machine management, including installation, configuration, and post setup troubleshooting
- Advise suppliers on hardware setup, driver configuration, BIOS and firmware settings, and network configuration for optimal performance
- Provide coverage for L1 support overflow during peak periods or incidents
- Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations
- Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead
- Collaborate with the engineering and support teams to flag and document systemic or recurring platform issues
You Are
- Experienced with Linux, especially Ubuntu, and comfortable troubleshooting from the command line
- Someone who enjoys debugging difficult problems and fixing broken systems
- Methodical and focused on finding root causes, not just temporary fixes
- Able to manage complex tickets independently
- A clear written communicator with an interest in AI infrastructure and GPU computing
Must-Haves
- Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions
- Proficiency with Docker, i …
More infrastructure roles at Vast.ai
-
GPU Systems Engineer – HPC / Parallel Computing
InfrastructureUS$200k–330k2mo
-
Senior Infrastructure Engineer
InfrastructureSeniorUS$200k–330k2mo
See also: Support Engineer jobs · AI jobs in Los Angeles · Vast.ai salaries · Python jobs · PyTorch jobs · TensorFlow jobs.
This listing is reproduced from Vast.ai's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All Vast.ai roles · AI salaries.