FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Cloud Systems Engineer – Site Reliability
TherapyNotes, LLC. Own and continuously improve use of Datadog across metrics, logs, traces, dashboards, monitors, alerts, and service-level views .
Posted 9/18/2026full-timeRemote • Pennsylvania • United StatesMid-LevelSenior💰 $110,000 - $150,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Systems Engineering and Cloud Engineering, with a strong focus on operational reliability, incident management, and automation. Proficient in using Datadog for observability and managing infrastructure as code with Terraform and Ansible.
Highest-signal resume keywords
Datadog ExpertiseInfrastructure As CodeIncident ManagementCloud-Based TechnologiesScripting Automation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Systems EngineeringCloud EngineeringDevOpsSRELinux SystemsNetworking FundamentalsScripting (Bash, PowerShell, Python)Infrastructure As Code (Terraform)Observability PlatformsProduction Systems Design
Soft Skills
CommunicationCollaborationProblem-Solving
Tools & Technologies
DatadogAzureKubernetesPrometheusGrafanaNew RelicAnsible
Certifications & Qualifications
BS Degree in Information SystemsBS Degree in Engineering
Industry Keywords
SaaS PlatformOperational ReadinessSLIsSLOsError BudgetsITSM PracticesHIPAA Compliance
Tech Stack
Tools & technologiesAnsibleAzureCloudDistributed SystemsGrafanaITSMKubernetesLinuxPrometheusPythonTerraform
About the role
Key responsibilities & impact- Own and continuously improve use of Datadog across metrics, logs, traces, dashboards, monitors, alerts, and service-level views
- Design, implement, and maintain high-availability, high-throughput, data- and compute-intensive critical systems for a 24×7 SaaS platform
- Partner with service owners to define and improve reliability through SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practices
- Participate in incident management as incident commander or technical responder, coordinating triage, restoration, escalation, communication, documentation, root cause analysis, and corrective actions
- Investigate issues across infrastructure and application layers using metrics, logs, distributed traces, and code-level context
- Improve deployment safety and service resilience through automated validation, recovery and rollback capabilities, reliability testing, and failure-mode analysis
- Ensure newly introduced systems are supportable and maintainable by development and operations
- Provide escalated technical guidance and support to technology teams
- Provide on-call coverage for production support and other duties as required
- Ensure systems and operational activities comply with organizational security, HIPAA, and operating policies
- Eliminate repetitive operational toil using Bash, PowerShell, Python, or Ansible
- Manage infrastructure as code using Terraform/OpenTofu and configuration automation using Ansible
Requirements
What you’ll need- BS degree in Information Systems, Engineering, or equivalent experience
- 5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SRE
- Experience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies; Azure and Kubernetes preferred
- Strong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems in production
- Expertise with an observability platform; Datadog experience strongly preferred
- Experience with Prometheus, Grafana, New Relic, or equivalent platforms is valuable
- Experience with scripting and operational automation using tools such as Bash, PowerShell, or Python, along with infrastructure-as-code and configuration-management practices
- Experience participating in production on-call rotations, incident response, root cause analysis, and post-incident improvement
- Experience working in Agile/DevOps environments and operating production services using ITSM practices where applicable
- Prior software development experience—or experience investigating application behavior through code, logs, and distributed traces—is a plus
Benefits
Comp & perks- Employer sponsored health, dental, vision, life, and disability insurance
- Retirement plan with company contribution
- Annual company profit sharing
- Personal development/training budget
- Open, collaborative work environment
- Extensive 2-week onboarding plan
- Comprehensive mentorship program