USI - Overall
Zhou et al. | Equally weighted mean of the four USI dimensions.
Kairos 1's benchmark results come from an early training checkpoint that beats Persimmon, Astra, Fable, and Gemini across 13 benchmarks.
We evaluate Kairos 1 with the User-Sim Index. Its behavioral measures compare generated conversation turns with human interactions using lexical features and rules. We report four shared dimensions: communication style, information patterns, clarification behavior, and reactions to errors.
Zhou et al. | Equally weighted mean of the four USI dimensions.
Zhou et al. | Measures how closely generated turns resemble human communication styles.
Zhou et al. | Measures alignment with the ways people share and organize information.
Zhou et al. | Compares clarification behavior in simulated and human interactions.
Zhou et al. | Compares how simulated users and people respond when an assistant makes errors.
Scores use a 0–100 scale. Higher scores indicate closer alignment with human behavior under the USI metric.
USI reference scores are drawn from published model reports, research papers, and reported evaluation results.
These tasks examine conversational realism, social interaction, and fidelity to individual users. We use the task definitions from the SOUL evaluation suite.
Naous et al. | Tests whether a model can generate realistic user messages grounded in a user's intent, using human conversation references.
Zhou et al. | Tests social interaction in challenging scenarios involving negotiation, collaboration, and conflict. Evaluation considers goals, relationships, and believability.
Hathidara et al. | Evaluates human-like user utterances across conversational tasks through lexical diversity and model-based judgments, separately from task success.
Dou et al. | Tests simulated user behavior during multi-turn document creation against annotated human–assistant conversations.
Dou et al. | Tests whether simulated learners reproduce human behavior in multi-turn math tutoring conversations.
Wu et al. | Tests whether a model reproduces an individual's book-related responses from their user history. Generated responses are evaluated by a model judge.
Wu et al. | Tests fidelity to an individual's email-writing behavior, comparing generated responses with real user responses through a model judge.
Wu et al. | Tests whether a model reproduces a user's political responses, conditioned on their history and evaluated against real responses.
Gemini 3.1 Pro reference scores come from its published results in the 𝒪dysSim report; the remaining scores come from the evaluation results used for this comparison.
We report all 17 comparisons. On these four, Kairos outperforms at least one reference model but does not lead.
Wu et al. | Tests whether generated opinions reflect a particular user's behavior and perspective, evaluated against human response references.
Kolluri et al. | Tests predictions of human responses in social science experiments. The SOUL evaluation reports correlation with human ratings.
Wu et al. | Tests fidelity to individual users' chat behavior, comparing simulated responses with their real conversation traces.
Li et al. | Tests personalized preference prediction: whether a model selects the response that best matches a user's represented preferences.
Scores use a 0–100 scale; higher is better within each task. Metrics differ across benchmarks, so compare scores within a chart.