Senior Software Engineer - BQL Reliability Engineering

USD 160,000-240,000 per year
SENIOR
✅ On-site

Tech Stack

Distributed Systems @ 4 Experimentation Java @ 4 Linux @ 7 Observability Python @ 4 Stress Testing

Details

As part of the BQL (Bloomberg Query Language) Reliability Engineering team, you will build software and platform capabilities that improve the reliability, resilience, and transparency of BQL and the services it depends on. You’ll work on engineering problems at significant scale, using software and automation to make reliability a built-in property of the platform rather than a purely operational concern.

Responsibilities

  • Design, build, and maintain software and self-service platform capabilities that enable engineering teams to understand, operate, and improve the reliability of BQL at scale.
  • Build tools and automated diagnostic capabilities that analyze telemetry and system behavior, helping engineers rapidly identify failures, regressions, and their root causes.
  • Develop software that improves incident detection and diagnosis, reducing Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) for high-severity incidents.
  • Engineer observability capabilities that turn metrics, logs, traces, and other system signals into actionable insights across BQL’s distributed architecture.
  • Partner with BQL engineering teams on system design, instrumentation, SLIs, and SLOs, ensuring reliability is built into services throughout the Software Development Lifecycle.
  • Improve platform resilience through engineering and experimentation, including load and stress testing, canary releases, controlled experiments, and failure testing.
  • Identify recurring operational problems and eliminate them through software, automation, and improvements to platform architecture.

Requirements

  • 4+ years of experience in Software Engineering, Reliability Engineering, Platform Engineering, or a related technical role.
  • Experience working with an object-oriented programming language such as C/C++, Python, or Java.
  • Strong knowledge of Linux/UNIX systems and experience developing or operating distributed applications in production.
  • Demonstrated experience improving the performance, availability, resilience, or scalability of mid- to large-scale systems.
  • Experience with production software delivery, including deployment, release management, testing, and safely introducing changes into distributed systems.
  • Strong analytical and problem-solving skills, including the ability to use production data and system telemetry to understand complex system behavior.
  • BA, BS, MS, or PhD in Computer Science, Engineering, or a related technical field.

Preferred Experience

  • Building developer platforms, reliability tooling, observability systems, or other infrastructure used by engineering teams.
  • Working with metrics, logs, distributed traces, time-series data, and other forms of production telemetry.
  • Defining and applying SLIs, SLOs, and other quantitative measures of system reliability.
  • Designing and executing load tests, stress tests, failure tests, canary releases, or other techniques for validating system resilience.
  • Applying statistical methods to understand system behavior and solve real-world engineering problems.
  • Querying and analyzing large-scale datasets in enterprise data environments.
  • Designing, executing, and analyzing A/B tests and other controlled experiments.
  • Operating in regulated or highly controlled environments.

Compensation and Benefits

Salary range: $160,000–$240,000 USD annually, plus benefits and bonus. Actual compensation may vary based on geographic location, work experience, market conditions, education or training, and skill level.

Benefits may include merit increases, incentive compensation for exempt roles, paid holidays, paid time off, medical, dental and vision coverage, short- and long-term disability benefits, 401(k) matching, life insurance, and wellness programs.

More jobs at Bloomberg

Similar jobs