CareerRiver

Senior Systems Engineer, OS Automation

CoreWeave, Inc. · San Francisco Bay Area

📍 Livingston, NJ / New York City, NY/ Sunnyvale, CA/ Bellevue, WA💰 $153,000 to $242,000via greenhousePosted 2026-07-27
Apply on company site ↗
CareerRiver pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to CoreWeave, Inc..
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at  www.coreweave.com . What You'll Do: HAVOCK owns the host software stack that turns a freshly provisioned bare-metal machine into a healthy CoreWeave Kubernetes worker — everything from power-on to a node joining the cluster: the OS image, boot-time configuration, and the services that decide which combination of software is valid for a given piece of hardware. As that problem space keeps growing, the only way to stay agile is to write software that manages the complexity, checks our work, and validates our assumptions — that's what this team builds. About the role: As a Senior Software Engineer on the Automation team, you will design, build, and operate the services, APIs, and libraries that sit behind our OS image, payload, and boot-configuration systems — the software platform other HAVOCK engineers and partner teams rely on to release, test, and ship node software quickly and safely. You'll work on a constraint-solver–based service that resolves compatibility between images, kernels, drivers, payloads, and hardware into a single validated configuration; an end-to-end test framework that validates OS images on real hardware; a library suite for declaratively configuring node storage; and natural-language tooling that lets stakeholders query and interact with our systems. This is a software- and platform-engineering role first, with a clear forward trajectory toward AI-assisted automation — log triage, regression detection, natural-language interfaces to infrastructure — but the core of the job is designing and shipping reliable services and APIs. Some of what you'll work on: Own and evolve a boot-configuration service that models complex compatibility and dependency relationships between OS images, kernels, drivers, payloads, and instance types as a constraint-solved graph, exposed through clean, well-specified interfaces. Extend our Kubernetes-native, end-to-end test framework that validates OS images and configuration on real hardware, plus the broader testing and validation story for the team. Build and maintain a library suite for declaratively configuring node storage — filesystems, mount options, block-device selection — with configurable strictness. Ship changes to our versioned, boot-time payload system (networking, storage, Kubernetes join) that's published as artifacts and consumed during node bring-up. Grow our natural-language / chat interface that lets stakeholders query and interact with the team's systems. Design and evolve versioned service contracts (gRPC / Connect-RPC, Protobuf) with strong correctness guarantees and robust validation. Build tooling that meaningfully shortens the build-and-release loop Lower the barrier to entry for everyone who touches this software, and lay the groundwork — clean interfaces, structured build/test metadata — for future AI-assisted automation across build triage, regression detection, and natural-language infrastructure tooling. Operate the services you build: participate in an on-call rotation for the team's services and own their reliability. Who You Are: 3+ years of professional software engineering experience building and operating backend services, platforms, or developer/infrastructure tooling. Strong proficiency in Go and/or Python , with the ability to work fluently across both. Experience designing and maintaining APIs and service contracts (REST, gRPC, or similar), with an eye for clean, well-specified, versioned interfaces. A demonstrated instinct for data modeling — representing relationships, constraints, and dependencies in code (graphs, constraint solving, relational models, or similar). Solid testing discipline: you write services that are testable, and you build the automation that proves they work. Comfort operating in a Kubernetes-based environment and reasoning about how software is built, packaged, deployed, and released. A working understanding of how Linux systems boot and are configured (the OS image / cloud-init / provisioning lifecycle), even if you haven't owned it end to end. A collaborative, software-development-lifecycle mindset (sprints, planning, code review, design docs) and the judgment to refactor toward simplicity. Preferred: Experience modeling complex problems in novel ways. gRPC / Connect-RPC and Protobuf experience, including evolving service contracts safely over time. Familiarity with bare-metal or node provisioning — PXE-style network boot, cloud-init, OS image building, firmware/driver enablement. Fluency with NVIDIA GPU platforms Rust experience and/or workflow orchestration tools like Argo Workflows. Linux packaging and repository management, configuration management (e.g., Ansible), and shell-based build pipelines. Interest in applying LLMs, RAG, and predictive modeling to large-scale infrastructure automation. Technical Stack: Languages: Go (primary), Python, bash/sh; Rust (test framework); Protobuf APIs & RPC: gRPC, Connect-RPC, HTTP/2, mTLS Orchestration & Infra: Kubernetes, Custom Resources, Helm, GitOps, workflow orchestration, Docker/containerd Node bring-up: cloud-init, OS image builds, GPU drivers, network-boot tooling Storage & Packaging: S3-compatible object storage, Linux package management, configuration management CI/CD & Observability: CI/CD pipelines, Prometheus-style metrics, dashboards The base salary range for this role is $153,000 to $242,000. The starting salary will be determined b

More San Francisco Bay Area jobs

San Francisco Bay Area jobs · Browse all locations