Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Ondo Finance

Site Reliability Engineer – Low-Latency Trading Systems

Ondo Finance

. Own production reliability for real-time trading services, including trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines .

Posted 9/18/2026full-timeRemote • United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in production reliability for real-time trading services, with strong capabilities in Kubernetes, AWS, and incident response. Proficient in programming with Go or Rust, and skilled in monitoring and alerting using Prometheus and Datadog.

Highest-signal resume keywords
Production Reliability EngineeringKubernetes ManagementAWS EKSGo ProgrammingIncident Response

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesAWSGoRustPromQLLinux InternalsNetworking FundamentalsPostgresS3Data Integrity
Soft Skills
Clear CommunicationIncident Management
Tools & Technologies
PrometheusDatadogGitOpsFluxSOPS
Industry Keywords
Real-Time TradingLatency-Sensitive SystemsMarket Data IngestionPnL ReconciliationOn-Call Rotation

Tech Stack

Tools & technologies
AWSFluxKubernetesLinuxPostgresPrometheusRustGo

About the role

Key responsibilities & impact
  • Own production reliability for real-time trading services, including trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines
  • Operate and evolve multi-region Kubernetes clusters on AWS EKS using GitOps with Flux and SOPS-encrypted secrets
  • Build and refine Prometheus metrics and alerting, Datadog logs and dashboards, and SLOs
  • Improve deployment safety through progressive rollouts, configuration reload behavior, and safeguards for live trading
  • Debug production incidents involving stale market data, exchange rate limits, WebSocket disconnects, order-lifecycle desynchronization, and trading-path latency regressions
  • Harden market data ingestion from Databento and venue-native REST/WebSocket feeds through staleness detection, failover, and replay
  • Build reconciliation and data-integrity tooling across live gauges, Postgres, and S3 Parquet data lake
  • Participate in on-call rotation covering US equity market hours and 24/7 crypto venues

Requirements

What you’ll need
  • 5+ years in SRE, production engineering, or infrastructure roles, with meaningful time supporting real-time or latency-sensitive systems
  • Strong programming ability in Go or Rust, and willingness to work in both
  • Deep, hands-on Kubernetes and AWS experience running stateful, latency-sensitive workloads in production
  • Fluent PromQL, structured-log analysis, and experience designing high-signal, low-noise alerts
  • Solid Linux internals and networking fundamentals
  • Ability to chase p99 regressions through the kernel, NIC, or GC
  • Experience participating in incident response and communicating clearly during and after incidents
  • Must be based in the United States

Benefits

Comp & perks
  • Competitive compensation including future token rights and/or equity according to preferences
  • Full medical, vision, and dental benefits
  • Flexible vacation policy (PTO)
  • Remote-first work arrangement
  • Opportunity to help shape the company’s vision, culture, and design practices
  • Collaboration with A+ colleagues and leading industry experts