Intern: AI Red Teaming (Fall 2026)

Realm Labs
Realm Labs

Software Engineering, Data Science

Sunnyvale, CA, USA

Posted on Aug 24, 2026
Intern: AI Red Teaming (Fall 2026)
Sunnyvale, CA
AI/ML
In office
Full-time

Role Overview

  • You will try to break the systems we build and the models we protect, and turn what you find into something the team can act on: a reproducible attack, an evaluation that catches it, and a written account of why it works.
  • Expect some mix of: eliciting unsafe behaviour from aligned LLMs and multi-modal models; prompt injection and tool-use abuse against agentic systems; automating attack generation and evaluation rather than hand-crafting one-off prompts; and measuring whether guardrails hold under pressure. Where RealmLabs' interpretability work gives you access to a model's internals, use it.
  • We aim for a paper or public technical report out of every internship, plus attacks that stay in our evaluation suite after you leave.

Expected Background: Adversarial ML and Red Teaming

  • Hands-on experience attacking or stress-testing models, from any direction: jailbreaks, prompt injection, adversarial examples, data poisoning, model extraction, or evaluating safety and moderation systems.
  • Able to read a paper and implement its attack.
  • (nice to have) Offensive security background outside ML: CTFs, vulnerability research, penetration testing.
  • (nice to have) Familiarity with agentic systems and their attack surface — tool calls, retrieval, memory, multi-agent orchestration.


Expected Background: ML

  • Machine learning tools: pytorch, huggingface, transformers, datasets.
  • Applied deep learning and LLM experience.
    • Training and evaluating deep models.
    • (nice to have) finetuning LLMs, multi-modal LLMs.
  • (nice to have) Familiarity with ML[NLP,LLM,Vision] interpretability methods, sparse autoencoders, linear probes — as a way of locating failure modes, not as an end in itself.

Expected Background: Software Engineering

  • Development environments and tools:
    • unix, git, basic clouds usage on AWS and/or GCP
    • jupyter
  • Programming:
    • python
    • (nice to have) “programming languages well-roundedness”
    • experience in statically-typed and functional languages

Compensation & Benefits

  • Market aligned compensation for interns in the bay area.

Requirements

  • Must be authorized to work in the USA or must be able to obtain CPT (Curricular Practical Training) approval from host university.
Ready to apply?
Powered by
First name *
Last name *
Email *
LinkedIn URL
Location *
Resume *
Click to upload or drag and drop here
Are you legally authorized to work in the United States? *
Earliest start date? *
Req ID: R14