Aaru
Head of Evaluations Research
NYC · Posted 2h ago
Job Description
About Aaru
Aaru builds simulations of human behavior. Each simulation contains a population of agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing, from product launches and policy changes to critical communications. Because the agents are simulated rather than recruited, they can reason through complex hypotheticals without fatigue or the response effects common in human studies.
The role
Evaluation Research determines whether Aaru's populations, predictions, and simulations correspond closely enough to the real world to support consequential decisions. The function defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities.
As Head of Evaluation Research, you will set Aaru's evaluation charter and lead the research needed to test its core systems. The strongest evaluations will be grounded in observed behavior and outcomes, including transactions, product usage, operational records, resolved events, and longitudinal decisions. You will work across Population Research, Prediction Research, Simulation Engineering, and customer-facing teams while preserving the independence needed to identify inconvenient results.
This is a hands-on research leadership role. You will design studies, construct evaluation datasets, write analysis code, inspect individual failures, and develop new measurements when existing benchmarks are inadequate. You will also recruit and lead a small team that can combine scientific rigor with a practical understanding of how research and engineering systems improve.
What you will do
Define a coherent evaluation agenda across population construction, predictive systems, agent behavior, group dynamics, and end-to-end simulations.
Turn broad questions about realism, accuracy, and usefulness into measurable constructs, decisive experiments, and clear decision criteria.
Build tests of population quality that assess whether each generated person forms a coherent whole, whether the population reproduces important relationships in the data, and whether rare but plausible profiles are represented.
Evaluate forecasts and other predictive outputs using temporal holdouts, prospective outcomes, calibration, ranking quality, subgroup performance, and the real cost of different errors.
Compare simulations with transactions, behavioral traces, product adoption, operational outcomes, market movements, and other records of what people actually did.
Design longitudinal and interaction-based evaluations that test how agents change over time, respond to new information, and influence one another.
Develop end-to-end studies that show whether better components lead to better answers for the decisions customers use Aaru to make.
Find failures hidden by aggregate metrics, especially those concentrated in important subgroups, rare cases, or changing environments.
Establish strong baselines, clean holdouts, contamination controls, and statistical standards appropriate to each research question.
Create diagnostic evaluations that help researchers identify why a system failed and whether a proposed fix generalizes.
Work with Simulation Engineering to make evaluations repeatable and versioned, and convert production outcomes and customer failures into durable test cases.
Produce evidence for customers and the public that is reproducible, appropriately scoped, and clear about uncertainty and limitations.
Communicate negative and inconclusive results with the same care as positive findings.
Hire, mentor, and lead exceptional evaluation researchers and research engineers while remaining a direct contributor.
Representative research directions
Construct a population using data available at one point in time, then test whether its future transactions, choices, or behavioral outcomes match what later occurred.
Develop methods for detecting contradictions, impossible combinations, unstable attributes, and unsupported specificity within individual profiles.
Recreate a historical decision environment using only information available before the outcome, then compare the simulation with the observed result.
Run prospective evaluations in which Aaru makes predictions before outcomes are known and tracks performance as those outcomes resolve.
Determine whether component measures such as row-wise coherence or overall accuracy predict the quality of an end-to-end simulation.
Compare agent-based simulation with direct forecasting, statistical models, and simpler segment-level approaches on the same outcome.
Measure whether simulated groups reproduce observed patterns in information diffusion, coordination, influence, or collective decision-making.
Study how population size, heterogeneity, interaction structure, model capability, and computation affect simulation fidelity.
Build tests for memorization, leakage, prompt sensitivity, unsupported certainty, and benchmark-specific overfitting.
Develop evaluation methods for settings where outcomes are delayed, noisy, only partially observed, or open to more than one reasonable interpretation.
How we work
A useful evaluation measures something consequential and helps the company learn. We compare systems with strong alternatives, protect held-out data, quantify uncertainty, and separate exploratory findings from evidence used to support a claim. Ground truth is often imperfect, so the quality and limits of the outcome data are part of the research problem.
Evaluation should make research faster by giving teams clear signals about what improved and what did not. It should also make Aaru more trustworthy by exposing failures early and keeping external claims aligned with the available evidence.
You might thrive in this role if
You have developed an original evaluation or measurement agenda in machine learning, behavioral science, computational social science, statistics, economics, psychometrics, or an environment with a comparable bar for rigor.
You have built evaluations that changed a research direction, model capability, product decision, or scientific conclusion.
You can define a difficult construct precisely enough to measure it without losing the underlying question.
You are comfortable with experimental design, observational data, sampling, uncertainty, statistical power, leakage, and condition shift.
You can write code, analyze large datasets, design studies, and inspect individual model failures.
You can work closely with the teams building a system while reaching independent conclusions about its quality.
You care more about an accurate result than a favorable one and are willing to revise your own evaluation when it proves inadequate.
You can explain technical evidence clearly to researchers, engineers, customers, company leadership, and the public.
You have led researchers or a major technical direction while remaining directly involved in the work.
You want to build in person, in New York, at high speed.
Strong candidates may also have
Work in ML evaluation, model behavior, forecasting, econometrics, psychometrics, causal inference, experimental economics, or measurement theory.
Experience evaluating LLM agents, multi-agent systems, synthetic populations, recommender systems, probabilistic models, or decision-support tools.
Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, or validation against operational outcomes.
Experience building evaluation platforms, regression suites, experiment-tracking systems, or shared research datasets.
A record of finding an important failure that standard metrics missed and developing a better way to measure it.
Experience communicating scientific or technical results in customer-facing, public, policy, or regulatory settings.
Success in this role looks like
Aaru has a clear and widely trusted definition of quality across populations, predictions, and end-to-end simulations.
Core evaluations are grounded in observed outcomes and reveal whether performance generalizes across time, domains, and groups.
Researchers receive measurements that are diagnostic enough to guide improvement while protected holdouts preserve the integrity of final results.
Important failures are found early, explained clearly, and converted into durable tests.
Product and company decisions rely on evidence that is reproducible, appropriately uncertain, and connected to real-world value.
Customers and the public can understand what Aaru has demonstrated, where the evidence is limited, and how confidence should be interpreted.
A small, exceptional team develops a reputation for evaluation work that is scientifically rigorous, practically useful, and unusually honest.