Kairos 1

Kairos 1's benchmark results come from an early training checkpoint that beats Persimmon, Astra, Fable, and Gemini across 13 benchmarks.

View benchmarksBenchmarks

User-Sim Index

We evaluate Kairos 1 with the User-Sim Index. Its behavioral measures compare generated conversation turns with human interactions using lexical features and rules. We report four shared dimensions: communication style, information patterns, clarification behavior, and reactions to errors.

Kairos 1Reference models

USI - Overall

Zhou et al. | Equally weighted mean of the four USI dimensions.

User-Sim Index0255075100Kairos 1: 86.11Kairos 186.11Persimmon: 77.82Persimmon77.82Grok 4.6: 70.38Grok 4.670.38GPT-6 Astra: 58.18GPT-6 Astra58.18Fable 5.1: 56.52Fable 5.156.52

USI - Communication

Zhou et al. | Measures how closely generated turns resemble human communication styles.

USI · Communication style0255075100Kairos 1: 82.11Kairos 182.11Persimmon: 66.02Persimmon66.02Grok 4.6: 58.50Grok 4.658.50GPT-6 Astra: 46.43GPT-6 Astra46.43Fable 5.1: 44.73Fable 5.144.73

USI - Information

Zhou et al. | Measures alignment with the ways people share and organize information.

USI · Information patterns0255075100Kairos 1: 94.96Kairos 194.96Persimmon: 91.78Persimmon91.78Grok 4.6: 88.00Grok 4.688.00Fable 5.1: 78.63Fable 5.178.63GPT-6 Astra: 73.52GPT-6 Astra73.52

USI - Clarification

Zhou et al. | Compares clarification behavior in simulated and human interactions.

USI · Clarification behavior0255075100Kairos 1: 83.52Kairos 183.52Grok 4.6: 82.90Grok 4.682.90Persimmon: 74.19Persimmon74.19GPT-6 Astra: 64.79GPT-6 Astra64.79Fable 5.1: 57.76Fable 5.157.76

USI - Error Reaction

Zhou et al. | Compares how simulated users and people respond when an assistant makes errors.

USI · Error reaction0255075100Kairos 1: 83.85Kairos 183.85Persimmon: 79.28Persimmon79.28Grok 4.6: 52.10Grok 4.652.10GPT-6 Astra: 47.97GPT-6 Astra47.97Fable 5.1: 44.97Fable 5.144.97

Scores use a 0–100 scale. Higher scores indicate closer alignment with human behavior under the USI metric.

USI reference scores are drawn from published model reports, research papers, and reported evaluation results.

Behavioral Benchmarks

These tasks examine conversational realism, social interaction, and fidelity to individual users. We use the task definitions from the SOUL evaluation suite.

UserLLM

Naous et al. | Tests whether a model can generate realistic user messages grounded in a user's intent, using human conversation references.

UserLLM0255075100Kairos 1: 91.71Kairos 191.71Fable 5.1: 68.50Fable 5.168.50Gemini 3.1 Pro: 67.70Gemini 3.1 Pro67.70GPT-6 Astra: 59.68GPT-6 Astra59.68

SOTOPIA-hard

Zhou et al. | Tests social interaction in challenging scenarios involving negotiation, collaboration, and conflict. Evaluation considers goals, relationships, and believability.

SOTOPIA-hard0255075100Kairos 1: 47.60Kairos 147.60Fable 5.1: 32.00Fable 5.132.00GPT-6 Astra: 30.49GPT-6 Astra30.49Gemini 3.1 Pro: 27.80Gemini 3.1 Pro27.80

MirrorBench

Hathidara et al. | Evaluates human-like user utterances across conversational tasks through lexical diversity and model-based judgments, separately from task success.

MirrorBench0255075100Kairos 1: 62.17Kairos 162.17Fable 5.1: 58.93Fable 5.158.93GPT-6 Astra: 57.00GPT-6 Astra57.00Gemini 3.1 Pro: 48.30Gemini 3.1 Pro48.30

SimArena - Document

Dou et al. | Tests simulated user behavior during multi-turn document creation against annotated human–assistant conversations.

SimArena · Doc0255075100Kairos 1: 84.72Kairos 184.72Fable 5.1: 83.10Fable 5.183.10GPT-6 Astra: 83.06GPT-6 Astra83.06Gemini 3.1 Pro: 83.00Gemini 3.1 Pro83.00

SimArena - Math

Dou et al. | Tests whether simulated learners reproduce human behavior in multi-turn math tutoring conversations.

SimArena · Math0255075100Kairos 1: 71.52Kairos 171.52Gemini 3.1 Pro: 71.50Gemini 3.1 Pro71.50Fable 5.1: 70.36Fable 5.170.36GPT-6 Astra: 69.17GPT-6 Astra69.17

Humanual - Book

Wu et al. | Tests whether a model reproduces an individual's book-related responses from their user history. Generated responses are evaluated by a model judge.

HuManuAL · Book0255075100Kairos 1: 64.73Kairos 164.73Fable 5.1: 64.27Fable 5.164.27Gemini 3.1 Pro: 62.40Gemini 3.1 Pro62.40GPT-6 Astra: 45.76GPT-6 Astra45.76

Humanual - Email

Wu et al. | Tests fidelity to an individual's email-writing behavior, comparing generated responses with real user responses through a model judge.

HuManuAL · Email0255075100Kairos 1: 51.24Kairos 151.24Fable 5.1: 50.80Fable 5.150.80GPT-6 Astra: 49.94GPT-6 Astra49.94Gemini 3.1 Pro: 46.90Gemini 3.1 Pro46.90

Humanual - Politics

Wu et al. | Tests whether a model reproduces a user's political responses, conditioned on their history and evaluated against real responses.

HuManuAL · Politics0255075100Kairos 1: 41.42Kairos 141.42Fable 5.1: 40.98Fable 5.140.98GPT-6 Astra: 36.48GPT-6 Astra36.48Gemini 3.1 Pro: 32.50Gemini 3.1 Pro32.50

Gemini 3.1 Pro reference scores come from its published results in the 𝒪dysSim report; the remaining scores come from the evaluation results used for this comparison.

Additional benchmark results

We report all 17 comparisons. On these four, Kairos outperforms at least one reference model but does not lead.

Humanual - Opinion

Wu et al. | Tests whether generated opinions reflect a particular user's behavior and perspective, evaluated against human response references.

HuManuAL · Opinion0255075100Fable 5.1: 43.23Fable 5.143.23Kairos 1: 42.28Kairos 142.28GPT-6 Astra: 40.86GPT-6 Astra40.86Gemini 3.1 Pro: 36.00Gemini 3.1 Pro36.00

SocSci210

Kolluri et al. | Tests predictions of human responses in social science experiments. The SOUL evaluation reports correlation with human ratings.

SocSci2100255075100Gemini 3.1 Pro: 78.00Gemini 3.1 Pro78.00Fable 5.1: 77.40Fable 5.177.40Kairos 1: 74.67Kairos 174.67GPT-6 Astra: 73.90GPT-6 Astra73.90

Humanual - Chat

Wu et al. | Tests fidelity to individual users' chat behavior, comparing simulated responses with their real conversation traces.

HuManuAL · Chat0255075100Fable 5.1: 27.30Fable 5.127.30Kairos 1: 24.33Kairos 124.33GPT-6 Astra: 21.92GPT-6 Astra21.92Gemini 3.1 Pro: 21.00Gemini 3.1 Pro21.00

AlignX

Li et al. | Tests personalized preference prediction: whether a model selects the response that best matches a user's represented preferences.

AlignX0255075100GPT-6 Astra: 76.60GPT-6 Astra76.60Gemini 3.1 Pro: 73.40Gemini 3.1 Pro73.40Kairos 1: 72.00Kairos 172.00Fable 5.1: 71.00Fable 5.171.00

Scores use a 0–100 scale; higher is better within each task. Metrics differ across benchmarks, so compare scores within a chart.