FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesAirflowAWSCloudDockerElasticSearchJavaScriptKafkaMongoDBNode.jsPostgresPythonRabbitMQSQLTypeScript
About the role
Key responsibilities & impact- Propose and build data-driven product features such as market and competitor intelligence, pricing benchmarks, contracting-body profiles, tender matching, alerting, and data-quality signals
- Own the data side of AI features
- Design and improve the retrieval layer behind Vera and the Dynamic RAG service, including document ingestion, conversion, chunking, embeddings, Qdrant indexing, Elasticsearch hybrid search, and reranking
- Use LLMs for structured extraction, classification, entity resolution, and summarisation
- Build evaluation sets to measure retrieval and answer quality
- Provide clean, structured, documented data and tools for Vera and the Tendios MCP server
- Define and measure data and retrieval quality, including source coverage, freshness, parsing accuracy, deduplication, and RAG answer quality
- Own and improve the collection pipeline, including Crawl Manager, Tenders Discovery, Collector, parsers, and RabbitMQ consumers
- Improve pipeline reliability, observability, and cost efficiency
- Maintain consistent read models in Elasticsearch and Qdrant
- Build the analytics layer on ClickHouse, orchestrated with Prefect and exposed through Metabase
- Replace manual Excel-based SaaS metrics reporting with a warehouse for MRR, churn, and NRR
- Write production Python for pipelines and data services
- Write TypeScript/Node.js in the NestJS Turborepo backend where data features touch the API
- Review code, set standards for data modelling and migrations, and document decisions in Confluence
- Contribute to data governance and security work related to ENS and ISMS compliance
- Report directly to the Head of Technology and collaborate with product, AI & Data, and backend squads
Requirements
What you’ll need- 6+ years in software engineering, with at least 3 years focused on data-intensive systems
- Strong Python and solid SQL
- Deep, hands-on PostgreSQL experience, including modelling, performance, and migrations at scale
- Comfortable reading and writing TypeScript/NestJS
- Experience building and running production data pipelines, including event-driven or queue-based architectures such as RabbitMQ or Kafka
- Experience with Elasticsearch/OpenSearch or ClickHouse
- Hands-on understanding of LLM applications, including RAG, embeddings, vector search, chunking strategies, prompt design, tool calling, and retrieval/answer quality evaluation
- Shipped at least one LLM or RAG feature to production
- Product mindset and ability to write clear proposals and defend them with product and business stakeholders
- Fluent Spanish and English
- Nice to have: web scraping and document parsing at scale, including PDFs and messy semi-structured sources
- Nice to have: Qdrant or other vector databases at scale, hybrid search, and reranking
- Nice to have: LLM observability and evaluation tooling such as Langfuse, or self-hosted open models such as Qwen on GPU infrastructure
- Nice to have: orchestration and ELT tooling such as Prefect, Airflow, dlt, or dbt
- Nice to have: large or legacy data migrations, especially MongoDB to PostgreSQL
- Nice to have: knowledge of public procurement, open data, or regulated environments
- Nice to have: Docker and cloud/hybrid infrastructure experience with Hetzner, DigitalOcean, or AWS
Benefits
Comp & perks- Full-time employment
- Work fully remote from anywhere in Spain, or hybrid from the Barcelona office
- Autonomous squads
- Modern technology stack
- Pragmatic culture
- Real ownership of data strategy
- Opportunity to work on interesting problems involving scraping at scale, entity resolution, polyglot persistence, and AI
- Direct impact on tender matching, Vera answers, and customer decisions
