Senior Hardware Engineer, Server Infrastructure

CoreWeave · New York, NY/ Bellevue, WA · Engineering

Posted 2026-09-30

Apply for this role →

About the Role

CoreWeave is seeking a highly skilled and motivated engineer to join our Hardware Engineering team. In this role, you will help design, develop, and optimize our server hardware infrastructure. You’ll collaborate closely with cross-functional teams, external vendors, and key stakeholders to deliver performant, reliable, and scalable hardware solutions that power CoreWeave’s rapidly growing infrastructure.

You will own server hardware from provisioning through decommission. This hands-on role combines engineering and operational support. Engineering work includes automation across the hardware lifecycle, hardware and firmware management services, monitoring and alerting, and qualification and bring-up of new platforms. Operational support includes acting as a senior point of contact for hardware escalations, driving deep root-cause analysis across hardware and firmware, and working quality and RMA issues through to resolution with server vendors and OEMs.

Both engineering and operational support are core responsibilities. You will support the systems you build and use what you learn from production failures to improve automation, telemetry, and platform design. You will work closely with data center operations teams, hardware technicians, and engineering teams to bring new regions online and keep existing infrastructure healthy.

What You’ll Do

Design and develop server hardware infrastructure to support CoreWeave’s high-performance workloads.

Automate all aspects of the server hardware lifecycle, from provisioning and configuration through firmware management, monitoring, and decommissioning.

Develop and maintain hardware and firmware management services that ensure reliability at scale.

Develop and implement monitoring and alerting for server hardware health, improving alert quality to support reliable on-call response.

Serve as a senior point of contact for hardware escalations, performing deep troubleshooting and root-cause analysis across hardware and firmware to drive long-term fixes.

Participate in an on-call rotation for hardware escalations and improve the runbooks, alerts, and tooling that make the rotation sustainable.

Collaborate with cross-functional teams to define hardware requirements, specifications, and system architecture.

Work with server vendors and OEMs to evaluate, qualify, and deploy new platforms and resolve firmware, quality, and RMA issues.

Support new data center region bring-up and hardware qualification.

Analyze hardware system performance, identify bottlenecks, and implement improvements to efficiency and resilience.

Establish and continuously refine processes for internal hardware testing, deployment, and performance optimization.

Create and maintain accurate documentation of hardware designs, specifications, test procedures, and results.

Support data center operations teams and hardware technicians with troubleshooting guidance, runbooks, and training so common issues can be resolved without engineering escalation.

Turn recurring production failures into automation, better telemetry, and improvements to platform design and vendor solutions.

Communicate status, trade-offs, and risks clearly to engineering, operations, and customer-facing stakeholders, including during active incidents.

Who You Are

Deep understanding of server hardware, components, and management technologies.

Proficiency in Ansible or Python, with hands-on experience programmatically interacting with server BMCs using Redfish or IPMI; Redfish preferred.

Experience collaborating with hardware vendors and OEMs to evaluate, qualify, and deploy server solutions.

Demonstrated experience supporting and troubleshooting production infrastructure, participating in on-call or escalation rotations, and driving incidents through to root-cause resolution.

Comfort with both building and automating systems and supporting infrastructure already in production.

Proven ability to stay current with technologies and trends in server and data center hardware.

Strong interest in automation and infrastructure scalability, with a commitment to continuous improvement.

Excellent technical documentation skills and attention to detail.

Strong analytical and problem-solving abilities, with a bias toward systematic, data-driven decisions.

Excellent written and verbal communication skills in English, with the ability to work effectively with technical teams and cross-functional stakeholders.

Preferred Qualifications

Experience bringing up new data center regions or standing up infrastructure in a new geography.

Experience with GPU platforms and rack-scale systems such as NVIDIA GB200 or GB300.

Experience with BMC technologies and management interfaces such as Redfish or IPMI at fleet scale.

Experience with firmware lifecycle management or hardware qualification programs in large-scale environments.

Strong Linux systems administration and debugging skills at fleet scale.

Familiarity with Kubernetes-based services, observability tools such as Prometheus and Grafana, and distributed production environments.

Experience designing alerting, runbooks, and self-service tooling that enable operations teams to resolve common hardware failures without engineering escalation.

Experience supporting external customers or partners in a technical troubleshooting capacity.

Wondering if You’re a Good Fit?

We believe in investing in our people and value candidates who bring diverse experiences to our teams—even if you aren’t a 100% skill or experience match. If some of this describes you, we’d love to talk.

You enjoy solving problems at the boundary of hardware, firmware, and software.

You want to both build systems and support them, and you use operational experience to improve what you build.

You enjoy tracing intermittent hardware failures to their root cause and validating that a fix addresses the underlying problem.

You look for opportunities to automate repetitive hardware procedures.

You work effectively with operations, engineering, and vendor teams to resolve complex technical issues.

You bring structure, accountability, and momentum to fast-moving environments.

Apply for this role →

← Back to all jobs