Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Peraton

Data Scientist

Peraton

. Build and curate the knowledge base for a generative AI/ML research effort .

Posted 9/29/2026full-timeRed Bank • New Jersey • United StatesMid-LevelSenior💰 $146,000 - $234,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing and curating knowledge bases for generative AI and machine learning, with strong capabilities in data extraction, validation, and analysis. Proficient in Python and familiar with deep learning frameworks, knowledge representation, and data versioning practices.

Highest-signal resume keywords
Python ProficiencyKnowledge Graphs ExperienceDeep Learning FrameworksData Versioning and ProvenanceStatistical Foundations

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Data Extraction PipelinesKnowledge GraphsMachine Learning ModelsStatistical AnalysisData ValidationSQLData SchemasMetadata StandardsAnnotation WorkflowsExperimental Design
Soft Skills
Clear DocumentationCollaboration
Tools & Technologies
NumPyPandasSciPyScikit-learnPyTorchSPARQLCypherAirflowPrefectNeo4j
Industry Keywords
Generative AIMachine LearningData ScienceData EngineeringOntologyRDFOWLProvenance TrackingVersion ControlScientific Datasets

Tech Stack

Tools & technologies
AirflowNeo4jNumpyPandasPythonPyTorchScikit-LearnSQL

About the role

Key responsibilities & impact
  • Build and curate the knowledge base for a generative AI/ML research effort
  • Develop data and knowledge extraction pipelines for machine learning models
  • Extract, structure, and validate design knowledge into knowledge graphs and ontologies
  • Engineer metadata, annotation, and provenance for program datasets
  • Analyze results, laboratory measurements, and simulation output to produce actionable findings
  • Design, build, populate, and maintain research knowledge bases under version control with provenance tracking
  • Develop rule-based, statistical, and ML-assisted pipelines that convert unstructured and semi-structured sources into structured, queryable knowledge
  • Establish quality metrics and validation procedures for extracted content
  • Engineer metadata, annotation schemas, and packaging for synthetic, simulated, and real datasets
  • Perform exploratory and inferential analysis on training data, simulation output, and evaluation results
  • Build dashboards and reports showing where generated waveforms succeed or fail against objectives
  • Feed findings back to AI model researchers and engineers
  • Contribute data and analysis sections to design reviews, monthly status reports, and dataset documentation
  • Coordinate with academic subcontractors on shared data and knowledge resources

Requirements

What you’ll need
  • Bachelor’s Degree or higher in Computer Science, Statistics, Mathematics, Electrical Engineering, or a related technical field
  • 5+ years of applied data science or data engineering experience, or MS with 3+ years
  • Strong Python proficiency including NumPy, pandas, SciPy, and scikit-learn
  • Experience with at least one deep learning framework, PyTorch preferred
  • Hands-on experience with knowledge graphs, ontologies, or structured knowledge representation, including RDF/OWL, Neo4j, or equivalent
  • Experience querying and validating knowledge graphs using SPARQL, Cypher, SHACL, or similar
  • Experience designing data schemas, metadata standards, and annotation workflows for large scientific or engineering datasets
  • Experience with data versioning and provenance
  • Solid statistical foundations including experimental design, hypothesis testing, and uncertainty quantification
  • Experience with SQL and at least one workflow or pipeline orchestration tool such as Airflow, Prefect, Dagster, or DVC
  • Ability to produce clear documentation including data dictionaries, dataset cards, and analysis reports
  • US Citizenship

Benefits

Comp & perks
  • Potential eligibility for overtime
  • Shift differential may be available
  • Discretionary bonus may be available