Arena (arena.ai) Jobs

Machine Learning Scientist - Open Source Lead

Arena (arena.ai)

Machine Learning Scientist - Open Source Lead

Reposted One Month Ago

Be an Early Applicant

Remote or Hybrid

Hiring Remotely in CA

Internship

Remote or Hybrid

Hiring Remotely in CA

Internship

The Machine Learning Scientist will lead open-source research, design experiments, develop novel methodologies, and analyze large-scale data to advance AI model evaluation and transparency.

The summary above was generated by AI

About Arena Intelligence

Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.

Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.

We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.

About the Role

Arena Intelligence is looking for a Machine Learning Scientist to lead our open-source research, including open data set and code releases, advancing how the world evaluates and understands AI models in the open. You’ll design, run, and share new methods and experiments that reveal what makes models useful, trustworthy, and capable, grounded in human preference signals and released openly for the full ecosystem and research community to build upon.

In this role, you’ll be responsible for taking our commitment to openness from principle to practice, curating high-impact datasets, developing new methodology and reproducible benchmarks, and releasing code that enables the research ecosystem to push AI evaluations forward. Your work will shape the public leaderboard, power community tools, and strengthen transparency in AI evaluation worldwide.

This role is deeply interdisciplinary, working with engineers, product teams, marketing, and the broader research community to advance how we compare models, analyze preference data, and understand factors like style, reasoning, and robustness. You’ll work closely with GTM teams as our spokesperson when it comes to outreach for our open research efforts: strengthening research partnerships, expanding research community participation, and championing programs that grow and support our research network.

If you’re excited by open-ended questions, rigorous evaluation, and scientific communication and outreach, you’ll find a meaningful home here. We’re looking for:

Hands-on experience training large-scale models, including reward models, preference models, and fine-tuning LLMs with methods like RLHF, DPO, and contrastive learning.
Strong foundation in ML and statistics, with a track record of designing novel training objectives, evaluation schemes, or statistical frameworks to improve model reliability and alignment.
Fluent in the full experimental stack, from dataset design and large-batch training to rigorous evaluation and ablation, with an eye for what scales to production.
Deeply collaborative mindset, working closely with engineers to productionize research insights and iterating with product teams to align research with user needs.
Comfortable being a visible representative of Arena Intelligence, engaging openly with the research community, and building a strong personal brand to help shape AI research culture.

You’ll

Design and conduct experiments to evaluate AI model behavior across reasoning, style, robustness, and user preference dimensions
Develop new metrics, methodologies, and evaluation protocols that go beyond traditional benchmarks
Analyze large-scale human voting and interaction data to uncover insights into model performance and user preferences
Communicate results with the broader research community via academic papers, educational content, conference talks
Collaborate with engineers to implement and scale research findings into production systems
Prototype and test research ideas rapidly, balancing rigor with iteration speed
Partner with model providers to shape evaluation questions and support responsible model testing
Contribute to the scientific integrity and transparency of the LMArena leaderboard and tools

You’ll have

PhD or equivalent research experience in Machine Learning, Natural Language Processing, Statistics, or a related field
Uses personal and professional platforms to amplify open research initiatives and invite collaboration.
Strong understanding of LLMs and modern deep learning architectures (e.g., Transformers, diffusion models, reinforcement learning with human feedback)
Proficiency in Python and ML research libraries such as PyTorch, JAX, or TensorFlow
Demonstrated ability to design and analyze experiments with statistical rigor
Experience publishing research or working on open-source projects in ML, NLP, or AI evaluation
Comfortable working with real-world usage data and designing metrics beyond standard benchmarks
Ability to translate research questions into practical systems and collaborate across engineering and product teams
Passion for open science, reproducibility, and community-driven research

Bonus Points

Skilled at public speaking, writing, and presenting research work to diverse audiences.
Actively participates in conferences, panels, and online forums to foster relationships and thought leadership.
Builds trust through transparent communication and consistent community engagement.
Serves as a go-to contact for external researchers, journalists, and partners.

What we offer

We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
The opportunity to work on cutting-edge AI with a small, mission-driven team
A culture that values transparency, trust, and community impact

Come help build the space where anyone can explore and help shape the future of AI.

Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.

Similar Jobs

StackAdapt

Scientist

22 Days Ago

In-Office or Remote

Mid level

AdTech • Marketing Tech

As an Applied Machine Learning Scientist, you will innovate and develop machine learning algorithms, write production code, and collaborate to implement and test these solutions based on historical data.

Top Skills: AlgorithmsData ScienceMachine LearningPython

Vertafore

Intern - Documentation Specialist

7 Hours Ago

Remote or Hybrid

18-18 Hourly

Internship

18-18 Hourly

Internship

Information Technology • Insurance • Software

Support review, rebranding, updating, and restructuring of English and French product documentation. Validate workflows, capture screenshots, apply adult learning principles, update release-driven changes, partner with product teams, and help establish repeatable documentation and version-control processes to improve usability and bilingual consistency.

Top Skills: Content Management SystemsDocumentation ToolsKnowledge BasesMS OfficeSharepoint

Dropbox

Director, Product Design

7 Hours Ago

Remote

217K-293K Annually

Expert/Leader

217K-293K Annually

Expert/Leader

Artificial Intelligence • Cloud • Consumer Web • Productivity • Software • App development • Data Privacy

Lead vision, strategy, and execution for Dropbox's next-generation design system and design technology platform. Build reusable patterns, components, tokens, and governance; partner with product, engineering, research, and brand; raise interaction, accessibility, and implementation quality; create adoption and measurement models; and develop a multidisciplinary team to scale coherent, high-quality multi-product experiences including AI-native capabilities.

Top Skills: Accessibility InfrastructureAi-Assisted ExperiencesComponent ArchitectureContent SystemsDesign SystemsDesign TechnologyDesign TokensFront-End EngineeringPrototyping Tools

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute