Senior Observability Engineer
Modernatx
via workday
Apply on company site ↗
CareerRiver pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Modernatx.
If you’re interested in this role, please apply in English and include an English version of your Resume/CV.
The Role:
Joining Moderna means advancing mRNA science to transform medicine. Work with exceptional global teams on a broad pipeline and build a career that makes a real difference for patients.
Moderna is strengthening its international business services hub in Warsaw, supporting our growing global operations. We welcome professionals ready to help advance our mission and shape the future of mRNA medicines.
This is an opportunity to define the future of enterprise observability by building a modern, AI-enabled observability platform that delivers scalable, resilient, and cost-efficient monitoring across Moderna's technology landscape. As a senior individual contributor, you will lead platform strategy, governance, and engineering while enabling intelligent operational insights across applications, infrastructure, cloud services, networks, containers, databases, and AI-powered systems. Working with open standards, automation, and emerging AI technologies, you will help improve reliability, operational excellence, and business outcomes across the region.
Here’s What You’ll Do
Own the enterprise observability platform, driving cost optimization, capacity planning, telemetry governance, and sustainable platform growth.
Manage and evolve Moderna's observability platform using technologies including OpenTelemetry, Prometheus, Grafana, VictoriaMetrics, and other modern observability solutions.
Lead platform governance, agent lifecycle management, architecture standards, roadmap development, and platform best practices.
Collaborate with technology vendors and open-source communities to influence product roadmaps and maximize platform value.
Design and build scalable, resilient, and cost-efficient observability architectures supporting applications, databases, hosts, containers, cloud services, networks, distributed systems, and AI/LLM-based workloads.
Develop and optimize telemetry pipelines for metrics, traces, and logs across hybrid cloud and on-premises environments.
Establish enterprise standards for monitoring, alerting, Service Level Objectives (SLOs), Service Level Indicators (SLIs), and proactive incident detection.
Enable self-service observability capabilities that accelerate troubleshooting, operational visibility, and platform reliability.
Design and implement observability capabilities for AI agents, LLM-powered applications, and agentic workflows, including monitoring prompts, responses, execution flows, latency, errors, taken consumption, and operational costs.
Define enterprise standards for AI observability, monitoring model performance, user interactions, reliability, and cost efficiency.
Implement AI observability solutions using Langfuse or similar technologies, enabling prompt analytics, optimization, intelligent alerting, debugging, and failure pattern detection.
Leverage AI and LLM technologies to support anomaly detection, operational insights, and root cause analysis.
Lead the enterprise logging strategy as a core pillar of the observability platform.
Design, build, and scale cost-efficient logging solutions using Grafana Loki and other modern open-source technologies.
Define enterprise logging standards, including centralized log ingestion, parsing, querying, retention policies, storage optimization, and query performance.
Partner with Security teams to ensure compliance, audit readiness, and forensic capabilities.
Integrate observability capabilities with incident management platforms such as PagerDuty to improve operational responsiveness.
Optimize on-call processes by ensuring alerts are meaningful, actionable, and routed effectively while supporting rapid incident resolution.
Provide real-time telemetry during incidents and contribute to root cause analysis activities.
Develop automation using Python, Terraform, Ansible, CI/CD pipelines, and infrastructure-as-code practices.
Implement self-healing capabilities and automated remediation to improve operational resilience.
Integrate the observability platform with enterprise technologies including ServiceNow, Jira, and other operational systems.
Develop dashboards and executive reporting that provide visibility into platform adoption, telemetry coverage, MTTA, MTTR, alert quality, reliability, operational performance, and cost efficiency.
Produce documentation, runbooks, knowledge articles, and technical training to support platform adoption and engineering consistency.
Participate in post-incident reviews and drive continuous improvement initiatives that strengthen operational excellence and foster a culture of observability and data-driven decision-making.
The key Moderna Mindsets you'll need to succeed in the role:
We act with dynamic range, driving strategy and execution at the same time at every step.
We obsess over learning. We don’t have to be the smartest, we have to learn the fastest.
Here’s What You’ll Need (Minimum Qualifications)
7+ years of experience in site reliability engineering, observability engineering, platform engineering, or related technical disciplines.
Strong hands-on experience designing, implementing, and operating modern observability platforms.
Experience with observability technologies such as OpenTelemetry, Prometheus, Grafana, VictoriaMetrics, Elastic, Datadog, Dynatrace, or similar solutions.
Strong understanding of metrics, logs, traces, telemetry pipelines, and SLO/SLI frameworks.
Experience supporting applications, infrastructure, containers, cloud services, and distributed systems.
Experience integrating observability platforms with incident management and operational workflows.
Hands-on experience with automation and infrastructure-as-code technologies such as Python, Terraform, Ansible, or Bash.
Experience supporting cloud-native and hybri
Browse all locations