Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
HighLevel

Staff Engineer – Distributed Systems

HighLevel

. Design, build, and scale backend systems and frontend interfaces powering the Threads & Composer experience inside HighLevel Conversations .

Posted 9/25/2026full-timeRemote • IndiaLeadWebsite

Tech Stack

Tools & technologies
CloudDistributed SystemsElasticSearchGoogle Cloud PlatformJavaScriptKubernetesMongoDBNode.jsNoSQLRedisSQLTypeScriptVue.jsGo

About the role

Key responsibilities & impact
  • Design, build, and scale backend systems and frontend interfaces powering the Threads & Composer experience inside HighLevel Conversations
  • Own backend and API design, data flows, performance, reliability, throughput, and correctness at scale
  • Own the Vue 3 UI layer end-to-end
  • Maintain architecture health of a billion-scale distributed system, including failure modes, capacity limits, consistency guarantees, and interactions across 50+ deployments
  • Approve critical-path designs and make architecture decisions across teams
  • Proactively identify and fix single points of failure, unbounded queues, missing idempotency, thundering herds, and data-loss windows
  • Prototype risky architectural changes, remediate serious incidents, and pair on complex cross-team bugs
  • Build resilience through degradation strategies, backpressure, isolation boundaries, and capacity models supporting 10%+ month-over-month growth
  • Raise engineering standards through design reviews, post-mortems, and reusable patterns
  • Establish safe AI-assisted engineering practices for critical systems
  • Work with Node.js/TypeScript, Go, GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch
  • Improve Workflows critical paths through ownership, capacity models, tested failure modes, and incident reduction

Requirements

What you’ll need
  • 10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems
  • Sole accountability for a production system through real failures
  • Deep command of queueing and asynchronous architectures, delivery semantics, ordering, backpressure, idempotency, exactly-once and at-least-once delivery
  • Strong knowledge of multiple SQL and NoSQL storage engines, consistency models, and indexing at scale
  • Expert-level depth in Redis or comparable in-memory systems, including failure modes under memory pressure and network partitions
  • Production experience with Kubernetes at scale, including resource limits, autoscaling behavior, and node-pool failures
  • Exceptional design communication through documentation, diagrams, and root-cause analyses
  • Fluency in Node.js and/or Go sufficient to prototype proposals and ship critical-path fixes
  • Comparable large-scale experience with hundreds of services or thousands of instances and billions of daily events strongly preferred
  • Experience with AI agents, GCP-native services, taking 0→1 systems to production, and hardening mature systems listed as bonus points

Benefits

Comp & perks
  • Global, remote-first organization
  • Equal Opportunity Employer
  • Voluntary demographic information process; information kept separate from application and not used in hiring decision
  • Privacy Policy available for review