We're an independent AI benchmarking company. We benchmark language models, inference providers, hardware, image/video and speech, and our Intelligence Index is widely cited when new models launch. Team of ~50, backed by Nat Friedman, Daniel Gross, Andrew Ng and others.
The people who build our evals also work directly with the AI labs, often on pre-release models. Flat structure, lots of ownership.
Main hiring needs:
\* Forward Deployed Engineer(FDE), Language Models: run our LLM benchmarking stack and work directly with the labs
\* Member of Technical Staff (MTS), Language Model Evaluations: build frontier evals (datasets, harnesses, benchmarks)
\* MTS, Inference: benchmark speed, quality and price across server-less inference providers
\* MTS, Hardware: benchmark GPUs, TPUs and custom silicon
\* MTS, Speech: own our TTS, STT and speech-to-speech evals
Also hiring for media generation, full stack, ML engineering, robotics and product roles.
Apply at [https://artificialanalysis.ai/careers](https://artificialanalysis.ai/careers).