← All jobs · Crusoe

Senior Manager, Engineering

Crusoe ·
51
AI-Agency
B62 U35
📍 Dublin, IE Manager 6+ yrs
KubernetesPrometheusVictoriaMetricsRabbitMQKafkaNATSTemporalGolangPythonTerraformAnsible
TL;DR

Senior Manager, Production Engineering at Crusoe leading a 24/7 operations team responsible for GPU infrastructure reliability, incident response, and observability. Reports to Director of Production Engineering with direct ownership of SLOs, monitoring, and team development across a scaling AI infrastructure platform.

Apply at Crusoe →
share:
you'll be redirected to the company's career page

Job description

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About This Role:

Crusoe is building the cloud infrastructure that powers the next generation of AI, and we're looking for a Senior Engineering Manager, Production Engineering to lead the team that keeps it running. This is a senior people management role reporting to the Director of Production Engineering — sitting at the intersection of deep technical leadership and organizational impact, with direct ownership over the reliability and operational health of Crusoe's production GPU infrastructure. You'll lead and develop a 24/7 team responsible for incident response, monitoring and alerting, automation, and continuous system improvement across a fast-scaling, high-stakes environment, while also shaping the broader strategy, culture, and structure of the function.

The ideal candidate is a seasoned technical leader who has built, scaled, and managed on-call operations teams in complex environments — someone who brings both rigor and vision to SLOs and postmortems, takes coaching and performance management seriously, and can drive alignment across engineering leadership on reliability strategy. If you're energized by the challenge of building a high-performing team while keeping complex systems reliable at scale, this role offers significant ownership and strategic impact at a critical moment in Crusoe's growth.

What You'll Be Working On:

What You'll Bring to the Team:

Bonus Points:

Benefits:

Crusoe also offers a competitive benefits package designed to support financial security, health, and overall well-being, including pension contributions, private health and dental insurance, income protection, life assurance and more.

Compensation:

Compensation will be paid as salary or hourly. Compensation to be determined by the applicant’s education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Apply at Crusoe →

More open roles at Crusoe

Crusoe · 🔄 synced 2h ago
Senior Director, AI Model LifeCycle
📍 San Francisco, US 💰 $301K–$355K · Director
Senior Director overseeing Crusoe's Model LifeCycle team, building ML infrastructure for fine-tuning, training, and aligning large language models at scale. Leads team of ML engineers on production LLM systems including SFT, PEFT, LoRA, RFT, and reinforcement learning pipelines.
PyTorchGolangPythonvLLMGPU systems
86
AI-core
Crusoe · 🔄 synced 2h ago
Staff Enterprise AI Automation Engineer
📍 San Francisco, US 💰 $190K–$230K 🛠 AI tools welcome at work · Staff
Staff Enterprise AI Automation Engineer at Crusoe designing and building agentic AI systems that orchestrate workflows across enterprise platforms. Focus on LLM integration, agent architecture, and scalable automation infrastructure.
PythonRESTGraphQLWorkatoAnthropic ClaudeGoogle Gemini
84
AI-core
Crusoe · 🔄 synced 2h ago
Senior Director of Engineering, Developer Experience
📍 San Francisco, US 💰 $301K–$355K 🛠 AI tools welcome at work · Director
Senior Director of Engineering for Developer Experience at Crusoe, an AI infrastructure company. Lead strategy and execution of internal developer platforms, CI/CD infrastructure, and AI-powered tooling to accelerate engineering velocity across the organization.
CICDDevOpsinternal APIsrepositoriesinfrastructure
76
AI-core
Crusoe · 🔄 synced 2h ago
Staff Product Manager, Managed Intelligence (SF/Sunnyvale)
📍 San Francisco, US 💰 $204K–$247K 🛠 AI tools welcome at work · Staff
Staff Product Manager at Crusoe leading product strategy for Managed Intelligence services. Focus on defining AI and agentic capabilities, model lifecycle, and scaling cloud products for AI-native companies.
PyTorchJAXTensorFlowKubernetesAWS
76
AI-core
Crusoe · 🔄 synced 2h ago
Senior Staff Software Engineer, AI Model LifeCycle
📍 San Francisco, US 💰 $237K–$318K · Staff
Senior Staff Software Engineer at Crusoe building managed platforms for AI model lifecycle, including fine-tuning systems, training pipelines, and reinforcement learning infrastructure for large language models.
PyTorchGolangPythonvLLMGPU systems
73
AI-fluent