FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Context Fusion AI Engineer – Autonomous Vehicles
NVIDIA. Design and develop learning-based, multimodal sensor-fusion systems that transform synchronized sensor history, ego-motion, navigation context, and driving context into a unified spatiotemporal world representation .
Posted 9/18/2026full-timeRemote • California • United StatesSenior💰 $184,000 - $356,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing multimodal sensor-fusion systems and deep-learning models for autonomous vehicles, with a strong focus on production-quality perception and state-estimation. Proficient in leveraging advanced architectures and techniques for real-time deployment and evaluation in complex environments.
Highest-signal resume keywords
Multimodal Sensor-Fusion DevelopmentDeep-Learning Model OptimizationC++ and Python ProgrammingTechnical Leadership in AV SystemsTransformer-Based Architectures
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Sensor-Fusion Systems3D PerceptionSemantic SegmentationObject Detection/TrackingEgo-Motion CompensationUncertainty EstimationDeep-Learning MethodsModel EvaluationDistributed TrainingCUDA Programming
Soft Skills
Problem-SolvingCollaborationTechnical Communication
Tools & Technologies
PyTorchNVIDIA GPUsTensorRTMixed-Precision TechniquesEdge Inference Toolchains
Industry Keywords
Autonomous VehiclesRoboticsMachine LearningComputer Vision3D Geometry
Tech Stack
Tools & technologiesCloudPythonPyTorchC++
About the role
Key responsibilities & impact- Design and develop learning-based, multimodal sensor-fusion systems that transform synchronized sensor history, ego-motion, navigation context, and driving context into a unified spatiotemporal world representation
- Build architectures that jointly reason over camera, LiDAR, radar, and vehicle-state inputs while handling calibration, synchronization, coordinate transforms, sensor latency, and uncertainty
- Develop end-to-end and multi-task models for road graph elements, semantic scene understanding, occupancy and free-space representations, including uncertain and occluded regions
- Develop scalable multimodal fusion architectures using Transformer-based early, late, and hierarchical fusion; BEV, point/voxel, and image-based representations; temporal context aggregation; and cross-modal attention
- Create training, fine-tuning, and evaluation pipelines for large-scale multimodal datasets
- Define multi-task objectives and metrics balancing perception quality, geometric consistency, prediction accuracy, latency, and safety-critical behavior
- Investigate foundation-model approaches for autonomous driving, including vision-language models, multimodal pre-training, representation learning, and efficient deployment of learned world models
- Collaborate with perception, mapping, prediction, planning, simulation, data, and embedded-software teams to turn research advances into production-quality AV systems
- Develop analysis and debugging tools for model failures, cross-sensor disagreement, long-tail scenarios, distribution shift, and regressions in closed-loop simulation and on-road evaluation
Requirements
What you’ll need- BS, MS, or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or a related technical field, or equivalent experience
- 8+ years of experience, including at least 2+ years in the AV or robotics industry and 2+ years of technical leadership experience
- Strong experience developing production-quality sensor-fusion, perception, state-estimation, or autonomous-driving systems
- Experience with learning-based multimodal perception or fusion involving two or more cameras, LiDAR, radar, map, navigation, and ego-motion signals
- Understanding of 3D geometry, coordinate frames, calibration, temporal synchronization, ego-motion compensation, tracking, uncertainty estimation, and sensor failure modes
- Experience with deep-learning methods for 3D perception, scene representation, occupancy/occlusion prediction, semantic segmentation, object detection/tracking, motion prediction, or planning
- Strong C++ and Python programming skills
- Hands-on experience developing, training, and optimizing deep-learning models in PyTorch
- Experience with Transformer, VLM, or multimodal foundation-model architectures, including pre-training, fine-tuning, distillation, quantization, or efficient inference
- Experience training and evaluating models at scale, including distributed training, dataset curation, offline evaluation, simulation-based validation, and production monitoring
- Ability to turn ambiguous AV problems into measurable technical objectives, build solutions, and drive them to deployment
- Experience with CUDA, distributed training, mixed-precision techniques, and efficient GPU inference using NVIDIA software and hardware is highly valued
- Experience with BEV, point-cloud/voxel, neural scene representation, 3D reconstruction, occupancy-flow, or spatiotemporal world-model methods is a plus
- Publications or open-source contributions in computer vision, robotics, machine learning, 3D perception, multimodal learning, or autonomous driving are a plus
- Experience optimizing models for automotive-grade real-time deployment using NVIDIA GPUs, TensorRT, CUDA, or edge inference toolchains is a plus
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score