← All jobs · Crusoe

Senior Staff Data Center Operations Engineer, GPU Hardware Architecture

Crusoe ·
59
AI-Agency
B62 U55
📍 San Francisco, US Staff 10+ yrs
PythonGoBashDCGMROCmNVIDIA HopperNVIDIA BlackwellAMD InstinctNVLinkInfiniBand
TL;DR

Senior Staff Data Center Operations Engineer at Crusoe building climate-aligned AI infrastructure. Bridges GPU hardware architecture with data center design, leads predictive maintenance strategies, and serves as technical authority on NVIDIA/AMD platforms for next-generation facilities.

Apply at Crusoe →
share:
you'll be redirected to the company's career page

Job description

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

The Mission

Crusoe is building the world’s most climate-aligned AI infrastructure. As we scale toward unprecedented power densities and liquid-cooled architectures, the gap between "Data Center Design" and "Silicon Reality" must be bridged.

We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical authority on GPU platforms within the Data Center Engineering and Operations organization. Your mission is twofold: act as the primary technical consultant to our Data Center Engineering team to ensure future facilities are built for next-gen silicon, and provide the Operations team with the specialized tooling, SOPs, and predictive strategies needed to maintain peak cluster health.

 

The Strategic Bridge

 

Key Responsibilities

 

Technical Requirements

 

Qualifications

Education: B.S. or M.S. in Electrical Engineering, Computer Engineering, or a related technical field.

 

Benefits:

 

Compensation Range

Compensation will be paid in the range of up to $179,000 -$218,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Apply at Crusoe →

More open roles at Crusoe

Crusoe · 🔄 synced 1h ago
Senior Director, AI Model Lifecycle
📍 San Francisco, US 💰 $301K–$355K · Director
Senior Director leading Crusoe's AI Model Lifecycle team, responsible for building fine-tuning systems, training pipelines, and managed platforms for large language models. Oversees team of ML engineers and manages end-to-end LLM training infrastructure including SFT, PEFT, LoRA, and reinforcement learning workflows.
PyTorchGolangPythonvLLMGPU systems
89
AI-core
Crusoe · 🔄 synced 1h ago
Staff Enterprise AI Automation Engineer
📍 San Francisco, US 💰 $190K–$230K 🛠 AI tools welcome at work · Staff
Staff Enterprise AI Automation Engineer at Crusoe designing and building agentic AI systems that orchestrate workflows across enterprise platforms. Focus on LLM integration, agent architecture, and scalable automation infrastructure.
PythonRESTGraphQLWorkatoAnthropic ClaudeGoogle Gemini
84
AI-core
Crusoe · 🔄 synced 1h ago
Staff Product Manager, Managed Intelligence (SF/Sunnyvale)
📍 San Francisco, US 💰 $204K–$247K 🛠 AI tools welcome at work · Staff
Staff Product Manager at Crusoe leading product strategy for Managed Intelligence services. Focus on defining AI and agentic capabilities, model lifecycle, and scaling cloud products for AI-native companies.
PyTorchJAXTensorFlowKubernetesAWS
76
AI-core
Crusoe · 🔄 synced 1h ago
Senior Staff Software Engineer, AI Model Lifecycle
📍 San Francisco, US 💰 $237K–$318K · Staff
Senior Staff Software Engineer at Crusoe building managed platforms for AI model lifecycle. Focus on fine-tuning systems, training pipelines, reinforcement learning, and dataset management for large language models at scale.
PyTorchGolangPythonvLLMGPU systems
73
AI-fluent
Crusoe · 🔄 synced 1h ago
Staff Software Engineer, AI Model Lifecycle
📍 San Francisco, US 💰 $208K–$279K · Staff
Staff Software Engineer at Crusoe building managed platforms for AI model lifecycle. Focus on fine-tuning systems, training pipelines, reinforcement learning, and dataset management for large language models at scale.
PyTorchGolangPythonvLLMGPU systems
73
AI-fluent