Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Communication @ 6
Debugging @ 4
GPU
LLM
Observability @ 4
OpenTelemetry @ 4
Profiling @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking a Staff Software Engineer to join its Observability team within the Infrastructure organization. The team owns monitoring and telemetry infrastructure used by engineers and researchers across metrics and logging pipelines, distributed tracing, profiling, error analytics, alerting, dashboards, and query interfaces.
The role focuses on building next-generation observability systems for large-scale GPU, TPU, and Trainium infrastructure, including high-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing and network visibility, and agentic diagnostic tools.
Responsibilities
- Design and build scalable telemetry ingest and storage pipelines for metrics, logs, traces, and error data across multi-cluster infrastructure.
- Build observability solutions that provide deep, low-overhead visibility into system behavior across the fleet.
- Own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth.
- Build instrumentation libraries, SDKs, and eBPF-based auto-instrumentation to emit high-quality telemetry, with and without code changes.
- Reduce mean time to detection and resolution through cross-signal correlation, from kernel-level events to application traces; unified query interfaces; and AI-assisted diagnostic tooling.
- Improve fleet-wide efficiency by turning continuous profiling and utilization telemetry into actionable optimization insights across CPU, memory, and accelerator fleets.
- Partner with Research, Inference, Product, and Infrastructure teams to ensure observability solutions meet the needs of each organization.
Requirements
- Hands-on experience building and operating large-scale observability or monitoring infrastructure.
- Deep, hands-on experience with observability signals end to end, from instrumentation through ingest to query and analysis.
- Understanding of high-throughput telemetry pipelines and the tradeoffs involved in collecting, storing, and querying operational data at scale.
- Comfort working below the application layer, including with the kernel, network stack, or hardware.
- Excellent communication skills and an interest in partnering with internal teams to improve operational visibility and incident response capabilities.
- Ability to build foundational infrastructure and navigate ambiguous, high-impact technical challenges independently and collaboratively.
- A bachelor's degree or equivalent combination of education, training, and experience. The field of study must be relevant to the role through coursework, training, or professional experience.
Preferred Qualifications
- 10+ years of relevant industry experience, including building and operating large-scale observability or monitoring infrastructure.
- Experience building or operating eBPF-based observability in production, including tracing, profiling, or network visibility.
- Experience running continuous profiling at fleet scale, including managing overhead budgets and symbolization.
- Kernel- and syscall-level debugging experience and performance engineering expertise.
- Experience profiling or instrumenting accelerator workloads.
- Experience operating metrics systems at very high cardinality or large-scale telemetry storage backends.
- Experience with OpenTelemetry instrumentation, collector pipelines, and tail-based sampling strategies.
- Interest in applying AI and LLMs to operational workflows such as automated root cause analysis, anomaly detection, or intelligent alerting.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration. Staff are currently expected to work from one of the company's offices at least 25% of the time, although some roles may require more office time.
Anthropic sponsors visas where possible and makes reasonable efforts to obtain visas for candidates who receive an offer, with support from an immigration lawyer.