🧠 What you will do: Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or measuring the wrong thing
🛠 What you need: Mid level profile, full-time rhythm, and hands-on with Data, Anyone AI Internal Team.
⚡ Why it is cool: Anyone Ai is hiring remote (fully remote - remote), so you can create real impact from anywhere.
📖 Full description is too long (BOOORING!!!).
👉 Read it on the official site, this is the quick preview version.