FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and maintaining AI-powered applications with a strong focus on reliability and observability. Proficient in Python development, BigQuery SQL, and implementing SLI/SLO frameworks to enhance application performance and user experience.
Highest-signal resume keywords
Python DevelopmentBigQuery SQL ExpertiseGCP ExperienceApplication-Level SLI/SLO FrameworksAgent-Based Systems
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Application ObservabilityAnomaly DetectionRoot Cause AnalysisFeature Flag ManagementAutomated RollbackComplex SQL QueriesSchema DesignDistributed TracingAgent Evaluation HarnessesOperational Analytics
Soft Skills
CollaborationLeadershipProblem Solving
Tools & Technologies
LookerBigQueryBigTableGKECloud LoggingCloud TraceCloud MonitoringLangGraph
Certifications & Qualifications
Bachelor's Degree in Computer ScienceMaster's Degree in Engineering
Industry Keywords
AI-Powered ApplicationsObservabilityProduction OperationsAIOpsEvent-Driven Systems
Tech Stack
Tools & technologiesBigQueryCloudGoogle Cloud PlatformKubernetesPythonSQL
About the role
Key responsibilities & impact- Own the reliability of AI-powered applications and features from the user's perspective
- Define, implement, and enforce feature-level SLIs, SLOs, and error budgets for APIs, RAG systems, AI agents, and user-facing applications
- Build and maintain application observability systems using Looker dashboards on BigQuery and BigTable
- Design and build LangGraph-based agents for automated issue identification and remediation
- Implement anomaly detection, root cause diagnosis, auto-rollback, feature-flag kill switches, and self-healing workflows
- Develop agent evaluation harnesses for benchmarking, multi-step workflow testing, non-deterministic outputs, and regression testing
- Write complex BigQuery SQL for usage trend analysis, anomaly detection, and operational analytics
- Design BigQuery table schemas optimized for observability and debugging
- Analyze application usage trends and adoption metrics to identify reliability risks, capacity needs, and degraded user experiences
- Partner with application development teams to embed deployment safety, structured logging, and distributed tracing practices
- Lead application-level incident response, root cause analysis, and blameless postmortems
- Build Python tooling and automation to reduce application-layer MTTD and MTTR
- Apply emerging AI techniques to improve platform reliability and developer productivity
- Collaborate with application developers, data engineers, infrastructure SREs, PMs, and leadership
Requirements
What you’ll need- 7+ years of experience in software engineering with significant focus on reliability, observability, or production operations
- Bachelor's or Master's Degree in Computer Science, Engineering, or a related technical discipline
- Strong Python development skills
- Experience building production tooling, automation, and agent-based systems
- Production GCP experience deploying and managing applications on GKE (Kubernetes)
- Deep SQL expertise with BigQuery, including complex queries, window functions, schema design, and cost optimization
- Hands-on experience with BigTable or equivalent high-throughput operational data technology
- Experience designing and operating application-level SLI/SLO frameworks, burn-rate alerting, and error budget policies
- Strong application-layer debugging skills, including distributed tracing, profiling, structured log analysis, and dependency mapping
- Experience building agent evaluation harnesses
- Familiarity with A2A protocols, streaming architectures, and event-driven systems
- Experience with feature flags, canary deployments, progressive rollouts, and automated rollback
- Experience with GCP observability services such as Cloud Logging, Cloud Trace, and Cloud Monitoring
- Exposure to AIOps concepts, including ML-driven anomaly detection, automated root cause analysis, and intelligent alerting
- Experience driving SLO adoption, postmortem processes, and reliability reviews
- Active engagement with the evolving AI ecosystem
- Hands-on GenAI application development experience with LangGraph, agent engineering, prompt design, and agentic workflows
- Experience building Looker dashboards and LookML models for operational observability
Benefits
Comp & perks- Medical, dental and vision insurance
- 401(k) plan with a Cisco matching contribution
- Paid parental leave
- Short- and long-term disability coverage
- Basic life insurance
- Cisco restricted stock unit grants may be available
- 10 paid holidays per full calendar year
- 1 floating holiday for non-exempt employees
- Paid birthday day off
- Paid year-end holiday shutdown
- 4 paid personal wellness days
- 16 days of paid vacation per full calendar year for non-exempt employees
- Flexible vacation time off program for exempt employees
- 80 hours of sick time off provided on hire date and each January 1st thereafter
- Up to 80 hours of unused sick time carried forward
- Additional paid time away for critical or emergency family issues
- Optional 10 paid volunteer days per full calendar year
- Annual bonuses for non-sales roles, subject to Cisco’s policies
- Incentive compensation for sales-plan employees, subject to applicable Cisco plan
