FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Site Reliability Engineer – Low-Latency Trading Systems
Ondo Finance. Own production reliability for real-time trading services, including trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in production reliability for real-time trading services, with strong capabilities in Kubernetes, AWS, and incident response. Proficient in programming with Go or Rust, and skilled in monitoring and alerting using Prometheus and Datadog.
Highest-signal resume keywords
Production Reliability EngineeringKubernetes ManagementAWS EKSGo ProgrammingIncident Response
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesAWSGoRustPromQLLinux InternalsNetworking FundamentalsPostgresS3Data Integrity
Soft Skills
Clear CommunicationIncident Management
Tools & Technologies
PrometheusDatadogGitOpsFluxSOPS
Industry Keywords
Real-Time TradingLatency-Sensitive SystemsMarket Data IngestionPnL ReconciliationOn-Call Rotation
Tech Stack
Tools & technologiesAWSFluxKubernetesLinuxPostgresPrometheusRustGo
About the role
Key responsibilities & impact- Own production reliability for real-time trading services, including trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines
- Operate and evolve multi-region Kubernetes clusters on AWS EKS using GitOps with Flux and SOPS-encrypted secrets
- Build and refine Prometheus metrics and alerting, Datadog logs and dashboards, and SLOs
- Improve deployment safety through progressive rollouts, configuration reload behavior, and safeguards for live trading
- Debug production incidents involving stale market data, exchange rate limits, WebSocket disconnects, order-lifecycle desynchronization, and trading-path latency regressions
- Harden market data ingestion from Databento and venue-native REST/WebSocket feeds through staleness detection, failover, and replay
- Build reconciliation and data-integrity tooling across live gauges, Postgres, and S3 Parquet data lake
- Participate in on-call rotation covering US equity market hours and 24/7 crypto venues
Requirements
What you’ll need- 5+ years in SRE, production engineering, or infrastructure roles, with meaningful time supporting real-time or latency-sensitive systems
- Strong programming ability in Go or Rust, and willingness to work in both
- Deep, hands-on Kubernetes and AWS experience running stateful, latency-sensitive workloads in production
- Fluent PromQL, structured-log analysis, and experience designing high-signal, low-noise alerts
- Solid Linux internals and networking fundamentals
- Ability to chase p99 regressions through the kernel, NIC, or GC
- Experience participating in incident response and communicating clearly during and after incidents
- Must be based in the United States
Benefits
Comp & perks- Competitive compensation including future token rights and/or equity according to preferences
- Full medical, vision, and dental benefits
- Flexible vacation policy (PTO)
- Remote-first work arrangement
- Opportunity to help shape the company’s vision, culture, and design practices
- Collaboration with A+ colleagues and leading industry experts