Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Cloudera

GraphRAG Engineer

Cloudera

. Provision, tune, and maintain production-grade Neo4j graph database and pgvector vector storage clusters .

Posted 9/29/2026full-timeRemote • Spain, Hungary, PolandMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in managing and optimizing Neo4j and pgvector databases, with a strong focus on hybrid search techniques and automated ingestion pipelines. Proficient in cloud environments (AWS, GCP) and Infrastructure-as-Code practices, ensuring efficient deployment and compliance in enterprise AI applications.

Highest-signal resume keywords
Neo4j Database ManagementPgvector ExpertiseHybrid Search TechniquesInfrastructure-as-Code (Terraform)Cloud Environments (AWS, GCP)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Cypher Query LanguageApache AvroKubernetes (EKS/GKE)DockerRelational Database Management (PostgreSQL)Graph Database ManagementPrompt OptimizationData ParsingAutomated Ingestion PipelinesVersion Control Systems
Tools & Technologies
LangChainLlamaIndexAWS MSKHashiCorp VaultOpenTelemetry (OTel)DatadogGrafana
Industry Keywords
Knowledge GraphsCI/CDZero-Trust SecurityDLP PII ScrubbingAgent Tool SpecificationSemantic Storage LayerGraph TraversalsContext EngineCode/Schema Lineage TrackingEnterprise AI

Tech Stack

Tools & technologies
ApacheAWSCloudDockerGoogle Cloud PlatformGrafanaKafkaKubernetesMicroservicesNeo4jPostgresSDLCSQLTerraformVault

About the role

Key responsibilities & impact
  • Provision, tune, and maintain production-grade Neo4j graph database and pgvector vector storage clusters
  • Engineer high-throughput index structures, cosine similarity vector indexes, and query optimizations for sub-second responses
  • Build automated ingestion pipelines parsing Git repositories, ASTs, Jira issue links, Apache Avro schemas, and CI/CD metadata into an enterprise knowledge graph
  • Connect distributed pipeline engines to hybrid retrievers combining SQL, Cypher graph traversals, and dense vector embeddings
  • Configure circuit breakers, confidence scoring thresholds, and step-limit constraints for autonomous agents
  • Integrate microservices and knowledge stores with the Enterprise AI Gateway
  • Maintain version-controlled system prompt structures in localized .ai/ spoke directories while following DLP PII scrubbing rules and token rate limits
  • Implement automated failover, backup restoration, and multi-cloud storage-tier cost controls across AWS and GCP
  • Own the semantic, vector, and graph storage layer powering the context engine for enterprise AI utilities and the Internal Developer Portal
  • Lead deployment of the SDLC Context Graph and GraphRAG Engine for CAB compliance, code/schema lineage tracking, and enterprise LLM proxy integrations

Requirements

What you’ll need
  • Deep operational and development experience with Neo4j (Cypher, APOC, causal clustering) or enterprise Knowledge Graphs
  • Proven expertise with pgvector (PostgreSQL), embeddings management, hybrid search techniques, and framework integrations (LangChain, LlamaIndex, or custom RAG pipelines)
  • Hands-on experience managing relational (PostgreSQL) and graph databases across AWS and GCP cloud environments
  • Proficiency in consuming Apache Avro payloads, streaming Kafka events (AWS MSK), and parsing structured/unstructured code and JSON artifacts
  • Practical understanding of Prompts-as-Code patterns, few-shot prompt optimization, and agent tool specification
  • Experience with Infrastructure-as-Code (Terraform) primitives, Kubernetes (EKS/GKE), Docker, and pull-based GitOps workflows
  • Exposure to HashiCorp Vault Transit encryption, OIDC keyless authentication, and zero-trust workload identities
  • Familiarity with OpenTelemetry (OTel) instrumentation for tracking vector search query latencies and LLM inference performance in Datadog or Grafana

Benefits

Comp & perks
  • Generous PTO Policy
  • Unplugged Days supporting work-life balance
  • Flexible WFH Policy
  • Mental & Physical Wellness programs
  • Phone and Internet Reimbursement program
  • Access to Continued Career Development
  • Comprehensive Benefits and Competitive Packages
  • Paid Volunteer Time
  • Employee Resource Groups