FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and building multimodal AI pipelines that integrate computer vision, speech-to-text, and language models, with a strong focus on programmatic extraction frameworks and data-centric AI methodologies. Proven ability to guide junior developers and establish best practices in engineering for data enrichment.
Highest-signal resume keywords
Applied AINatural Language ProcessingComputer VisionPythonPyTorch/TensorFlow
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Multimodal Pipeline DevelopmentProgrammatic Extraction FrameworksFine-Tuning Open-Source ModelsData-Centric AIActive Learning
Soft Skills
MentoringCollaboration
Tools & Technologies
WhisperDSPyInstructorPydanticVLLM
Industry Keywords
Social Media Data PipelinesRAG ArchitecturesVector DatabasesHigh-Throughput Inference EnginesSynthetic Data Generation
Tech Stack
Tools & technologiesPythonPyTorchTensorflow
About the role
Key responsibilities & impact- Design and build multimodal AI pipelines combining computer vision, speech-to-text (Whisper), and LLMs into automated ingestion workflows
- Implement programmatic extraction frameworks (Instructor, Pydantic, DSPy, Outlines) to enforce strict, validated JSON outputs from generative models
- Transition workloads from commercial APIs to fine-tuned open-source models (Llama 3/4, Qwen, DeepSeek) running on local inference engines (vLLM, TGI, Groq)
- Deploy data-centric AI frameworks (Cleanlab, Snorkel, Active Learning) to automate millions of labels and clean noisy datasets
- Guide junior developers, establish engineering best practices, and set architectural roadmaps for data enrichment
- Convert raw, unstructured social content (short-form video, audio, captions, and visual metadata) into high-precision structured data products
Requirements
What you’ll need- 5+ years of experience in Applied AI, Natural Language Processing, or Computer Vision
- Proven experience building multimodal pipelines integrating vision, speech, and text models
- Mastery of Python, PyTorch/TensorFlow, and production LLM orchestration tools (DSPy, Instructor, Pydantic, LangChain/LlamaIndex)
- Hands-on experience fine-tuning and serving open-source models using vLLM, TGI, or similar high-throughput inference engines
- Experience with vector databases (Qdrant, pgvector, Pinecone) and RAG architectures at scale
- Preferred: Prior experience with social media data pipelines (TikTok, Instagram Reels, YouTube Shorts)
- Preferred: Background in data-centric AI, active learning, or synthetic data generation
Benefits
Comp & perks- Remote work arrangement
