Embodied AI / VLA Research Engineer

Maxinsights
Maxinsights

Software Engineering, Data Science · Full-time

Santa Clara, CA, USA

Posted on Oct 9, 2026
Job Description:

Position Overview

In this role, you will work on the research, training, optimization, and real-world deployment of embodied AI models, including Vision-Language-Action (VLA) models, World Models, and related robotics foundation models.

You will work across multimodal perception, robot control, long-horizon task planning, large-scale robot data, and model deployment. This is a highly hands-on role that involves both model development and real-world robotic system integration.

Key ResponsibilitiesEmbodied AI Model Development

  • Research and develop Vision-Language-Action (VLA), World Model, and other embodied AI foundation models.
  • Work on areas including:
    • Robotic manipulation
    • Multimodal perception and control
    • Long-horizon task planning
    • Vision-language-action reasoning
    • Robot-environment interaction
  • Develop and optimize models for real-world robotic applications.
  • Translate research ideas into practical models and systems that can operate reliably on physical robots.
Foundation Model Training & Optimization

  • Train and optimize embodied AI foundation models using large-scale real-robot datasets and egocentric human demonstration data.
  • Design approaches for incorporating multimodal inputs such as:
    • Vision
    • Force
    • Tactile sensing
    • Proprioception
    • Other robot and environmental signals
  • Develop and evaluate multimodal fusion architectures.
  • Conduct model training, evaluation, benchmarking, and performance optimization.
  • Analyze model performance and identify opportunities to improve training efficiency, generalization, and real-world performance.
Robotics Data Pipeline

  • Build and improve large-scale robotics data pipelines for model training.
  • Develop processes for:
    • Data cleaning
    • Data filtering
    • Resampling
    • Data augmentation
    • Data quality evaluation
  • Design scalable data pipelines capable of supporting large volumes of robot and human demonstration data.
  • Work closely with data and robotics teams to improve dataset quality and training efficiency.
Model Deployment & Real-Robot Testing

  • Deploy trained models to physical robot systems.
  • Perform real-world robot debugging, testing, and performance optimization.
  • Diagnose issues across models, sensors, software, and robotic hardware.
  • Iterate between model training and real-world testing to improve system performance.
  • Help ensure models operate reliably and consistently in real-world environments.

QualificationsRequired

  • 1+ years of relevant industry or research experience in machine learning, robotics, computer vision, embodied AI, or a related field.
  • Strong understanding of deep learning and modern machine learning methods.
  • Experience with PyTorch or similar deep learning frameworks.
  • Experience training and evaluating machine learning models.
  • Strong programming skills in Python and familiarity with relevant ML/robotics tooling.
  • Understanding of multimodal learning, computer vision, robotics, or related areas.
  • Ability to work in a fast-paced startup environment and take ownership of technical problems from research through implementation.
  • Strong problem-solving and debugging skills.

Strong Plus

Real-World Robotics Deployment

  • Experience deploying and debugging machine learning models on physical robot systems.
  • Ability to bring models from development into real-world robotic environments.
  • Experience troubleshooting and stabilizing robotic systems in production or experimental environments.

Multimodal / VLA Models

  • Experience working with force, tactile, or other multimodal sensing.
  • Experience designing or training multimodal fusion models.
  • Hands-on experience with VLA models, including model design, training, evaluation, or real-world applications.
  • Experience with robotic manipulation or embodied AI systems.

Large-Scale Distributed Training

  • Experience with large-scale distributed model training.
  • Familiarity with DDP, DeepSpeed, FSDP, or similar distributed training frameworks.
  • Experience optimizing training performance, GPU utilization, memory usage, or training throughput.
  • Experience working with large-scale datasets and distributed data pipelines.

Ideal Candidate

We are looking for an engineer who is excited about the intersection of foundation models and physical robotics.

The ideal candidate is:

  • Hands-on and comfortable moving between research, coding, experimentation, and real-world robot testing.
  • Interested in solving problems that cannot be addressed through simulation or software alone.
  • Comfortable working with large-scale datasets and modern foundation-model architectures.
  • Able to take ownership of a problem from data → training → evaluation → deployment → real-world iteration.
  • Comfortable working in an early-stage environment where priorities can move quickly.
  • Curious about emerging VLA, World Model, and embodied AI research and able to translate new ideas into working systems.

Why Join MaxInsights?

  • Work directly on embodied AI and robotics foundation models.
  • Work with large-scale real-world robotics and human demonstration data.
  • Gain hands-on experience across the full AI development lifecycle, from data pipelines to real-robot deployment.
  • Work in a fast-moving startup environment with significant ownership and technical autonomy.
  • Collaborate with teams working at the forefront of robotics and foundation-model development.

Default Benefits:

  • Health insurance
  • Vision care
  • Dental coverage
  • 401(k)
  • Paid holidays
  • PTO (Paid Time Off)
  • Sick leave