Startups

Zibra Labs

Founding ML Researcher (Post Training)

← All Open Roles

Founding ML Researcher (Post Training) San Francisco

In person

Post Training

Team

Zibra Labs is building the post-training and inference infrastructure that currently only exists at the frontier labs. Our vision is to enable everyone to post-train and serve open weight models at a fraction of the cost with higher capability. The research side of that vision owns the recipes: how an open weight model gets post-trained cheaply and served cheaply without giving up capability. You will sit next to the team building the runtime, close enough that a recipe needing a new primitive can get one.

What You Will Work On

  • Post-training recipes for open weight models that hold up outside the benchmark they were tuned against.
  • Reinforcement learning at the algorithm level: reward design, rollout efficiency, and keeping long-horizon runs stable.
  • Making a served model cheaper without making it worse, through quantization, distillation, sparsity, and speculative decoding.
  • Evaluation that survives contact with reality, catching the regressions a leaderboard number hides.
  • Shaping the product direction: what we post-train, what we serve, and what ships next are research calls as much as engineering ones.

About You

  • You have trained models at a scale where the infrastructure fought back, and you can tell which failures were the recipe and which were the machine.
  • You read the literature closely enough to know which results reproduce.
  • You have taste in experiment design: you can pick the smallest run that answers the question.
  • You write code other people can run. Research is not an excuse.
  • You're high agency: you can take a capability goal, turn it into experiments, and ship the result.
  • You care about the quality of your work.

Jargon That Might Be Useful

None of this is a checklist. It is a map of the territory, so you can tell whether it is the territory you want to be in.

  • Post-training. SFT, DPO, GRPO, PPO
  • RL. veRL, SkyRL
  • Training. PyTorch, JAX, FSDP, DeepSpeed
  • Inference. SGLang, vLLM, TensorRT
  • Efficiency. FP8 and INT4 quantization, distillation, speculative decoding, LoRA
  • Evaluation. lm-evaluation-harness, task-specific harnesses

Apply

Send a note about what you have built and what you want to work on, with anything that shows the work: a paper, a repo, a training run you debugged to the bottom.

Email Us →

Our address: careers (at) zibralabs (dot) ai

Sourced 2026-09-25