Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
AWISEE

Lead AI Systems Architect, Data-Centric AI, Multimodal

AWISEE

. Design and build multimodal AI pipelines combining computer vision, speech-to-text (Whisper), and LLMs into automated ingestion workflows .

Posted 10/9/2026full-timeRemote • SwedenSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and building multimodal AI pipelines that integrate computer vision, speech-to-text, and language models, with a strong focus on programmatic extraction frameworks and data-centric AI methodologies. Proven ability to guide junior developers and establish best practices in engineering for data enrichment.

Highest-signal resume keywords
Applied AINatural Language ProcessingComputer VisionPythonPyTorch/TensorFlow

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Multimodal Pipeline DevelopmentProgrammatic Extraction FrameworksFine-Tuning Open-Source ModelsData-Centric AIActive Learning
Soft Skills
MentoringCollaboration
Tools & Technologies
WhisperDSPyInstructorPydanticVLLM
Industry Keywords
Social Media Data PipelinesRAG ArchitecturesVector DatabasesHigh-Throughput Inference EnginesSynthetic Data Generation

Tech Stack

Tools & technologies
PythonPyTorchTensorflow

About the role

Key responsibilities & impact
  • Design and build multimodal AI pipelines combining computer vision, speech-to-text (Whisper), and LLMs into automated ingestion workflows
  • Implement programmatic extraction frameworks (Instructor, Pydantic, DSPy, Outlines) to enforce strict, validated JSON outputs from generative models
  • Transition workloads from commercial APIs to fine-tuned open-source models (Llama 3/4, Qwen, DeepSeek) running on local inference engines (vLLM, TGI, Groq)
  • Deploy data-centric AI frameworks (Cleanlab, Snorkel, Active Learning) to automate millions of labels and clean noisy datasets
  • Guide junior developers, establish engineering best practices, and set architectural roadmaps for data enrichment
  • Convert raw, unstructured social content (short-form video, audio, captions, and visual metadata) into high-precision structured data products

Requirements

What you’ll need
  • 5+ years of experience in Applied AI, Natural Language Processing, or Computer Vision
  • Proven experience building multimodal pipelines integrating vision, speech, and text models
  • Mastery of Python, PyTorch/TensorFlow, and production LLM orchestration tools (DSPy, Instructor, Pydantic, LangChain/LlamaIndex)
  • Hands-on experience fine-tuning and serving open-source models using vLLM, TGI, or similar high-throughput inference engines
  • Experience with vector databases (Qdrant, pgvector, Pinecone) and RAG architectures at scale
  • Preferred: Prior experience with social media data pipelines (TikTok, Instagram Reels, YouTube Shorts)
  • Preferred: Background in data-centric AI, active learning, or synthetic data generation

Benefits

Comp & perks
  • Remote work arrangement