LP
Leading Path Consulting
Data Engineer

TS/SCI w/Poly - AI/ML Data Engineer

On-siteSeniorData Engineerposted 1mo ago
✦Role summaryAI-generated

The AI/ML Data Engineer will design and maintain the data infrastructure needed to train custom language translation models, including pipelines for text and audio data. They will collaborate with data scientists to enable low‑resource language training while ensuring secure handling of classified information under TS/SCI clearance.

Skills required

About this role

Al/ML Engineer

Multiple locations available including Warrenton, Chantilly and McLean, VA! You choose!

TS/SCI w/Poly required

The Data Engineer will work closely with the team to advance Human Language Technologies (HLT), with a focus on refining text and audio translation capabilities. This position will be responsible for establishing an environment to train custom models for the top language types currently available in our triage tool. Additionally, this role will enable the organization to develop the capacity to train low-resource languages which may emerge as future priorities. The individual in this role will work closely with Data Scientists, providing comprehensive support and ensuring seamless coverage. A professional who utilizes statistical analysis, programming skills, and machine learning techniques to collect, clean, analyze, and interpret large datasets, extracting valuable insights and creating predictive models

KEY RESPONSIBILITIES

 Develop, fine-tune, evaluate, and optimize multilingual machine translation models (e.g., NLLB, Opus-MT, MarianMT) to improve translation quality for low-resource languages.

 Build, preprocess, and manage multilingual text and speech datasets for model training, evaluation, and continuous improvement.

 Design, develop, and maintain scalable data pipelines and end-to-end MLOps workflows for data ingestion, model training, deployment, monitoring, and lifecycle management.

 Develop and deploy cloud-native machine learning solutions using AWS services such as SageMaker, Step Functions, and Bedrock.

 Deploy and support machine learning models in production using containerized environments and CI/CD best practices.

 Engineer features and optimize datasets to improve machine learning model performance.

 Conduct testing, validation, benchmarking, and troubleshooting of machine learning models and data pipelines.

 Research and evaluate emerging AI, machine learning, NLP, and speech technologies for mission applications.

Requirements

EDUCATION AND EXPERIENCE

 Bachelor’s Degree in Computer Science, Electrical or Computer Engineering or a related technical discipline, or the equivalent combination of education, technical training, or work/military experience

 10+ years of related software engineering experience.

REQUIRED QUALIFICATIONS

 Experience developing software applications using Python.

 Experience training, fine-tuning, evaluating, and optimizing machine learning and deep learning models using modern frameworks and best practices.

 Familiarity with DevOps and MLOps principles, including CI/CD, infrastructure automation, model lifecycle management, monitoring, version control, and software delivery best practices.

 Hands-on experience with AWS cloud services and AI/ML offerings, including S3, EC2, IAM, VPC, SageMaker, Bedrock, Lambda, and related services.

 Experience developing and deploying containerized applications using Docker.

DESIRED QUALIFICATIONS

 Experience fine-tuning transformer-based language translation models (e.g., NLLB, Opus-MT, MarianMT) and working with Hugging Face Transformers.

 Familiarity with experiment tracking, model registries, and dataset versioning tools such as MLflow, Weights & Biases, or DVC.

 Experience with distributed training frameworks such as Ray or PyTorch Distributed

 Hands-on experience with machine learning frameworks such as PyTorch or TensorFlow.

 Experience designing, implementing, and maintaining production-grade MLOps pipelines and automated machine learning workflows supporting model training, deployment, monitoring, and lifecycle management using technologies such as AWS Step Functions, Apache NiFi, Apache Airflow, or similar orchestration platforms.

 Experience developing multilingual NLP, speech processing, or language translation solutions.

 Familiarity with large language models (LLMs), transformer architectures, and generative AI technologies.

 Familiarity with NLP frameworks such as spaCy, NLTK, Stanford CoreNLP, or similar libraries.

 Strong analytical, problem-solving, and communication skills.

 Ability to work independently and collaboratively in a multidisciplinary environment.

 Familiarity with audio processing frameworks such as Librosa, PyAudioAnalysis, OpenSMILE, or similar technologies.

Benefits

· Generous starting salary with annual raises

· Fully paid Medical, Dental and Vision coverage premiums for employee and all their immediate dependents

· Generous PTO and Comp time

· 11 paid holidays

· 6% 401(k) auto contribution

· Company-funded life insurance

· Tuition and Training expense reimbursements

· Financial rewards for employee referrals

· Cash bonus SPOT Awards for efforts above and beyond your role like contributing to proposals, working billable overtime, receiving customer recognition or awards, supporting BD/marketing material creation and more!

✕ position closed

This role is no longer accepting applications. It’s kept here for reference — check out the similar open roles below.

Similar open roles

D
NEW

Director, Data Platform Engineering

Domino's·Ann Arbor, MI, United States
On-siteSeniorData Engineer
yesterday
V
NEW

Data Engineer

Visa·IN - Bengaluru, India
On-siteData Engineer
yesterday
L
NEW

Junior Data Engineer

Littelfuse·Kaunas - Donelaicio
On-siteJuniorData Engineer
yesterday
D
NEW

VP, Senior Cloud Engineer, Data Platform, Group Technology

Dbs·Singapore - East
On-siteSeniorData Engineer
yesterday
C
NEW

Data Engineer

Citigroup·Pune Maharashtra India
On-siteMidData Engineer
yesterday