Skip to content

SimulationEvals

SimulationEvals runs AI-driven end-to-end conversation simulations against your CXAS agent. Instead of scripting exact utterances, you describe goals and success criteria — and a Gemini model figures out what to say at each turn to try to achieve them. This is a great way to test how your agent handles realistic, messy, unpredictable conversations.

Here are the key concepts:

  • Step (Pydantic model) — a single goal within a simulation, with a goal, success_criteria, optional response_guide, and a max_turns limit. Steps can also include a static_utterance for when you want a fixed first message, and inject_variables for seeding session state.
  • StepStatus enum — tracks whether each step is NOT_STARTED, IN_PROGRESS, or COMPLETED.
  • simulate_conversation() — drives the full multi-turn loop, returning an LLMUserConversation object that contains the transcript, step progress, and expectation results.
  • generate_report() — produces a SimulationReport with two DataFrames: goal progress and expectation results. It renders as styled HTML in a Jupyter notebook.

Quick Example

from cxas_scrapi import SimulationEvals
from cxas_scrapi.utils.rate_limiter import RateLimiter

app_name = "projects/my-project/locations/us/apps/my-app-id"

# Optional: configure a rate limiter to pace simulation turns and prevent quota exhaustion
limiter = RateLimiter(requests_per_minute=30.0)
sim = SimulationEvals(app_name=app_name, rate_limiter=limiter)

test_case = {
    "steps": [
        {
            "goal": "User wants to check their account balance",
            "success_criteria": "Agent provides a numeric balance and account status",
            "max_turns": 5,
        },
        {
            "goal": "User asks to dispute a charge",
            "success_criteria": "Agent acknowledges the dispute and provides a reference number",
            "max_turns": 8,
        },
    ],
    "expectations": [
        "The agent should never ask for the full credit card number",
        "The agent should offer to escalate if it cannot resolve the dispute",
    ],
}

# Run the simulation
conversation = sim.simulate_conversation(
    test_case=test_case,
    console_logging=True,
)

# View the report
report = conversation.generate_report()
print(report)  # Colorized in terminal, styled HTML in Jupyter

Naturalness Metric (optional)

SimulationEvals can also grade how human the agent sounds. The metric is opt-in: a test case that does not declare a naturalness_metric block behaves exactly as before, with no extra Gemini call and no extra keys in the results. See the Local Simulations guide for the YAML schema and worked examples.

  • NaturalnessConfig (Pydantic model) — the per-test-case configuration: which model grades, which qualities are scored, the label bands, the turn/conversation blend weight, the perceived-latency bands and weight, and an optional pass_threshold. Unknown keys are tolerated, so newer YAML still loads.
  • NaturalnessResult — the aggregate result for one simulation: overall_score, overall_label, per-turn gradings, conversation-level factors, factor_averages, label_counts, latency_ms_by_turn, failure_reasons, and passed. Exposed as conversation.naturalness_result.
  • TurnNaturalness — the grading for a single agent turn: its label, 1–5 score, justification, and the factors behind it.
  • NaturalnessFactor — one scored quality (e.g. pacing) with its 1–5 score, an optional short value descriptor, and the evidence-citing reason.
  • NaturalnessLabel enum — Bot-like, Transitional, or Human-like.
  • evaluate_naturalness() — grades a simulation trace directly, optionally against the agent's recorded audio and the platform's per-turn perceived latency. Returns None (after logging) when there is nothing to grade or the grading call fails, so the metric can never break a run.

In audio modality the agent's recordings are attached to the grading call automatically, so pacing and emotion are judged from what the caller heard. perceivedLatency is measured from the conversation trace rather than judged, and a turn past latency_failure_threshold_ms fails the simulation outright.

Reference

SimulationEvals

SimulationEvals(app_name, rate_limiter=None, expectations_only=False, deployment_id=None, vertex_location='global', naturalness=None, **kwargs)

Bases: Apps

Wrapper class to simulate entire multi-turn conversations with a CXAS Agent.

Source code in src/cxas_scrapi/evals/simulation_evals.py
def __init__(
    self,
    app_name: str,
    rate_limiter: RateLimiter | None = None,
    expectations_only: bool = False,
    deployment_id: str | None = None,
    vertex_location: str = "global",
    naturalness: bool | dict[str, Any] | None = None,
    **kwargs: typing.Any,
) -> None:
    self.expectations_only = expectations_only
    self.vertex_location = vertex_location
    # Run-level override for the optional Naturalness Metric. `None`
    # defers entirely to each test case; True enables it everywhere with
    # defaults; a dict is merged over the test case's own block.
    self.naturalness = naturalness
    project_id = app_name.split("/")[1]
    location = app_name.split("/")[3]
    super().__init__(project_id=project_id, location=location, **kwargs)
    # Must follow super().__init__: `Common` assigns `self.app_name`
    # unconditionally, and it is not handed the app name here, so setting
    # this any earlier would be silently overwritten with `None`.
    self.app_name = app_name
    self.sessions_client = Sessions(
        app_name,
        deployment_id=deployment_id,
        rate_limiter=rate_limiter,
        **kwargs,
    )
    self.tools_map = Tools(app_name=app_name, **kwargs).get_tools_map()

    self.genai_client = GeminiGenerate(
        project_id=self.project_id,
        location=self.vertex_location,
        credentials=self.creds,
    )

simulate_conversation

simulate_conversation(test_case, sim_user_model=_DEFAULT_GEMINI_MODEL, eval_model=_DEFAULT_GEMINI_MODEL, session_id=None, console_logging=True, modality='text', capture_agent_audio=False, background_noise_file=None, burst_noise_files=None, use_tool_fakes=False, voice_config=None, initial_utterance=_FIRST_UTTERANCE, skip_playback_wait=False, single_bidi_stream=False, max_turns=None, naturalness=None, **kwargs)

Runs the simulated conversation loop.

Parameters:

Name Type Description Default
test_case dict[str, Any]

The test case dictionary defining evaluation steps.

required
sim_user_model str | None

The Gemini model used for the simulated user.

_DEFAULT_GEMINI_MODEL
eval_model str | None

The Gemini model used for evaluating expectations.

_DEFAULT_GEMINI_MODEL
console_logging bool

Whether to print interaction transcript to the console.

True
single_bidi_stream bool

For audio modality, keep one persistent bidi WebSocket open for the whole conversation instead of opening a new connection per turn (the default).

False
max_turns int | None

Maximum number of conversation turns. Defaults to the test_case's max_turns setting, or 30 if unspecified.

None
naturalness bool | dict[str, Any] | None

Overrides the test case's naturalness_metric block. None (the default) defers to the test case, so simulations that never declare the metric are unaffected.

None
Source code in src/cxas_scrapi/evals/simulation_evals.py
@cleanup_session_dir
def simulate_conversation(
    self,
    test_case: dict[str, Any],
    sim_user_model: str | None = _DEFAULT_GEMINI_MODEL,
    eval_model: str | None = _DEFAULT_GEMINI_MODEL,
    session_id: str | None = None,
    console_logging: bool = True,
    modality: str = "text",
    capture_agent_audio: bool = False,
    background_noise_file: str | None = None,
    burst_noise_files: list[str] | None = None,
    use_tool_fakes: bool = False,
    voice_config: dict[str, Any] | None = None,
    initial_utterance: str = _FIRST_UTTERANCE,
    skip_playback_wait: bool = False,
    single_bidi_stream: bool = False,
    max_turns: int | None = None,
    naturalness: bool | dict[str, Any] | None = None,
    **kwargs: Any,
) -> LLMUserConversation:
    """Runs the simulated conversation loop.

    Args:
        test_case: The test case dictionary defining evaluation steps.
        sim_user_model: The Gemini model used for the simulated user.
        eval_model: The Gemini model used for evaluating expectations.
        console_logging: Whether to print interaction transcript to
            the console.
        single_bidi_stream: For audio modality, keep one persistent
            bidi WebSocket open for the whole conversation instead of
            opening a new connection per turn (the default).
        max_turns: Maximum number of conversation turns. Defaults to
            the test_case's max_turns setting, or 30 if unspecified.
        naturalness: Overrides the test case's `naturalness_metric`
            block. `None` (the default) defers to the test case, so
            simulations that never declare the metric are unaffected.
    """
    sim_user_model = sim_user_model or _DEFAULT_GEMINI_MODEL
    eval_model = eval_model or _DEFAULT_GEMINI_MODEL
    if session_id is None:
        session_id = str(uuid.uuid4())
    naturalness_config = parse_naturalness_config(
        test_case,
        naturalness if naturalness is not None else self.naturalness,
    )
    voice_config = voice_config or test_case.get("voice_config")
    eval_conv = LLMUserConversation(
        genai_client=self.genai_client,
        genai_model=sim_user_model,
        test_case=test_case,
        max_turns=max_turns,
        initial_utterance=initial_utterance,
    )

    # Initialize audio paths tracking
    eval_conv.agent_audio_paths = {}
    current_sim_turn = 0

    interactive_session = None
    if modality == "audio" and single_bidi_stream:
        client = self.sessions_client
        interactive_session = client.create_interactive_session(
            session_id=session_id,
            capture_agent_audio=capture_agent_audio,
            background_noise_file=background_noise_file,
            use_tool_fakes=use_tool_fakes,
            skip_playback_wait=skip_playback_wait,
            voice_config=voice_config,
        )
        interactive_session.start()

    try:
        if console_logging:
            print(
                f"Starting simulated conversation with session ID: "
                f"{session_id}"
            )

        # Initialize the first turn manually
        user_utterance, variables = eval_conv.next_user_utterance()
        accumulated_variables = {}
        if variables:
            accumulated_variables.update(variables)

        detailed_trace = []
        detailed_trace.append(f"User: {user_utterance}")

        while user_utterance:
            if modality == "audio" and interactive_session:
                response = interactive_session.send_turn(
                    user_utterance,
                    accumulated_variables,
                )
                # Check if session ended via WebSocket endSession
                if isinstance(response, dict) and response.get(
                    "session_ended"
                ):
                    if response.get("connection_error"):
                        err_msg = (
                            f"Interactive session WebSocket error: "
                            f"{response['connection_error']}"
                        )
                        raise BidiSessionError(err_msg)
                    break
            else:
                response = self._send_request_with_retry(
                    session_id=session_id,
                    user_utterance=user_utterance,
                    variables=accumulated_variables,
                    modality=modality,
                    console_logging=console_logging,
                    turn_num=current_sim_turn,
                    capture_agent_audio=capture_agent_audio,
                    background_noise_file=background_noise_file,
                    burst_noise_files=burst_noise_files,
                    use_tool_fakes=use_tool_fakes,
                    voice_config=voice_config,
                )
            if not response:
                break

            # Extract and save the agent turn audio WAV if present
            # in response.
            if response and getattr(response, "agent_audio_paths", None):
                audio_path = response.agent_audio_paths.get(0)
                if audio_path:
                    paths = eval_conv.agent_audio_paths
                    paths[current_sim_turn] = audio_path

            if console_logging:
                self.sessions_client.parse_result(response)

            agent_text, trace_chunks, session_ended, tool_calls = (
                self._parse_agent_response(response)
            )
            detailed_trace.append("\n".join(trace_chunks))

            if session_ended:
                if agent_text:
                    eval_conv._add_agent_response(agent_text)
                eval_conv._add_agent_tool_calls(tool_calls)
                # Ensure the final agent response is evaluated
                # so that steps_progress is updated on session end.
                eval_conv._next_user_utterance()
                if console_logging:
                    print(
                        "\nSession has been closed by the Agent via "
                        "end_session tool."
                    )
                # Mark current step as completed if the session ending
                # is a valid success (escalation evals)
                for prog in eval_conv.steps_progress:
                    criteria = prog.step.success_criteria.lower()
                    if prog.status != StepStatus.COMPLETED and (
                        "escalat" in criteria
                        or "transfer" in criteria
                        or "being transferred" in criteria
                    ):
                        prog.status = StepStatus.COMPLETED
                        prog.justification = (
                            "Agent ended session via escalation/transfer — "
                            "matches success criteria."
                        )
                break

            # Get the next simulated user utterance based on the agent's
            # response
            eval_conv._add_agent_tool_calls(tool_calls)
            user_utterance, variables = eval_conv.next_user_utterance(
                agent_text
            )
            if variables:
                accumulated_variables.update(variables)
            if user_utterance:
                detailed_trace.append(f"User: {user_utterance}")

            current_sim_turn += 1

        if console_logging:
            self._print_completion_status(eval_conv)

        self._evaluate_expectations(
            eval_conv,
            detailed_trace,
            eval_model,
            console_logging,
            capture_agent_audio=capture_agent_audio,
        )
        self._evaluate_naturalness(
            eval_conv,
            detailed_trace,
            eval_model,
            console_logging,
            naturalness_config,
            session_id=session_id,
            modality=modality,
        )
        eval_conv._session_id = session_id
        eval_conv.session_id = session_id
        eval_conv._detailed_trace = detailed_trace
        eval_conv.detailed_trace = detailed_trace
        return eval_conv
    finally:
        if interactive_session:
            interactive_session.close()

export_results_to_golden

export_results_to_golden(results, output_path=None)

Exports simulation results to a Golden Evaluation YAML file.

Fetches the full conversation trace for each simulation from the platform to ensure accuracy.

Parameters:

Name Type Description Default
results list[dict[str, Any]]

The list of results returned by run_simulations.

required
output_path str | None

Optional local path to save the generated YAML.

None

Returns:

Type Description
str

The generated YAML string.

Source code in src/cxas_scrapi/evals/simulation_evals.py
def export_results_to_golden(
    self,
    results: list[dict[str, Any]],
    output_path: str | None = None,
) -> str:
    """Exports simulation results to a Golden Evaluation YAML file.

    Fetches the full conversation trace for each simulation from the
    platform to ensure accuracy.

    Args:
        results: The list of results returned by run_simulations.
        output_path: Optional local path to save the generated YAML.

    Returns:
        The generated YAML string.
    """
    conversations_list = []

    for res in results:
        turns = self._get_turns(res)
        if not turns:
            continue

        expectations = [
            e["expectation"] for e in res.get("expectation_details", [])
        ]
        params = res.get("session_parameters", {})

        conversations_list.append(
            GoldenConversation(
                conversation=res.get("name", "Simulated_Conversation"),
                turns=turns,
                expectations=expectations,
                session_parameters=params,
            )
        )

    dataset = GoldenConversations(conversations=conversations_list)
    yaml_content = yaml.dump(
        dataset.model_dump(exclude_none=True),
        sort_keys=False,
        allow_unicode=True,
    )

    if output_path:
        with open(output_path, "w", encoding="utf-8") as f:
            f.write(yaml_content)

    return yaml_content

Step

Bases: BaseModel

StepStatus

Bases: str, Enum

SimulationReport

SimulationReport(goals_df, expectations_df=None, naturalness_df=None, naturalness_headline='')

A report containing both Goals and Expectations DataFrames.

Source code in src/cxas_scrapi/evals/simulation_evals.py
def __init__(
    self,
    goals_df: pd.DataFrame,
    expectations_df: pd.DataFrame | None = None,
    naturalness_df: pd.DataFrame | None = None,
    naturalness_headline: str = "",
) -> None:
    self.goals_df = goals_df
    self.expectations_df = expectations_df
    self.naturalness_df = naturalness_df
    self.naturalness_headline = naturalness_headline

NaturalnessConfig

Bases: BaseModel

Per-test-case configuration for the Naturalness Metric.

Unknown keys are preserved rather than rejected so that a YAML file written against a newer version of the library still loads here.

NaturalnessResult

Bases: BaseModel

Aggregated Naturalness Metric result for a whole simulation.

TurnNaturalness

Bases: BaseModel

Naturalness grading for one agent turn.

NaturalnessFactor

Bases: BaseModel

A single scored conversational quality, e.g. pacing or grammarStyle.

NaturalnessLabel

Bases: str, Enum

Graded label assigned to an agent turn and to the whole call.

evaluate_naturalness

evaluate_naturalness(gemini_client, model_name, trace, config, audio_paths=None, latency_ms_by_turn=None)

Grades the naturalness of every agent turn in a simulation trace.

Parameters:

Name Type Description Default
gemini_client Any

A GeminiGenerate instance.

required
model_name str

Fallback model, used when the config does not pin one.

required
trace list[str]

The simulation's detailed_trace.

required
config NaturalnessConfig

Resolved metric configuration.

required
audio_paths dict[int, str] | None

Optional map of turn index to the agent's recorded audio, as either a gs:// URI or a local WAV path. Used only when config.use_audio is set.

None
latency_ms_by_turn dict[int, float] | None

Optional map of turn index to the platform's reported perceived latency in ms. When supplied, each turn gains a computed perceivedLatency factor.

None

Returns:

Name Type Description
A NaturalnessResult | None

class:NaturalnessResult, or None when there was nothing to

NaturalnessResult | None

grade or the grading call failed. Failures are logged rather than

NaturalnessResult | None

raised so the metric can never break a simulation run.

Source code in src/cxas_scrapi/evals/naturalness.py
def evaluate_naturalness(
    gemini_client: Any,
    model_name: str,
    trace: list[str],
    config: NaturalnessConfig,
    audio_paths: dict[int, str] | None = None,
    latency_ms_by_turn: dict[int, float] | None = None,
) -> NaturalnessResult | None:
    """Grades the naturalness of every agent turn in a simulation trace.

    Args:
        gemini_client: A ``GeminiGenerate`` instance.
        model_name: Fallback model, used when the config does not pin one.
        trace: The simulation's ``detailed_trace``.
        config: Resolved metric configuration.
        audio_paths: Optional map of turn index to the agent's recorded audio,
            as either a `gs://` URI or a local WAV path. Used only when
            ``config.use_audio`` is set.
        latency_ms_by_turn: Optional map of turn index to the platform's
            reported perceived latency in ms. When supplied, each turn gains a
            computed ``perceivedLatency`` factor.

    Returns:
        A :class:`NaturalnessResult`, or ``None`` when there was nothing to
        grade or the grading call failed. Failures are logged rather than
        raised so the metric can never break a simulation run.
    """
    turns = extract_agent_turns(trace)
    if not turns:
        logger.info("Naturalness metric skipped: no agent turns in trace.")
        return None

    target_model = config.model or model_name
    prompt: Any = _build_prompt(turns, config)

    if config.use_audio and audio_paths:
        prompt = _build_audio_contents(prompt, turns, audio_paths)

    try:
        output: NaturalnessOutput | None = gemini_client.generate(
            prompt=prompt,
            model_name=target_model,
            response_mime_type="application/json",
            response_schema=NaturalnessOutput,
        )
    except Exception as exc:  # noqa: BLE001 - metric must never break a run
        logger.error("Naturalness grading failed: %s", exc)
        return None

    if not output or not output.turns:
        logger.warning("Naturalness grading returned no turn gradings.")
        return None

    latencies = latency_ms_by_turn or {}
    # Injected before scoring so the weighted turn mean includes latency.
    _attach_latency_factors(
        output.turns,
        latencies,
        config,
        speech_index_by_turn={t.turn_index: t.speech_index for t in turns},
    )

    graded_turns = _normalise_turns(output.turns, turns, config)
    for factor in output.conversation_factors:
        factor.score = int(min(max(factor.score, 1), 5))

    overall = _aggregate(graded_turns, output.conversation_factors, config)
    label_counts: dict[str, int] = {
        label.value: 0 for label in NaturalnessLabel
    }
    for turn in graded_turns:
        label_counts[turn.label.value] += 1

    passed = None
    if config.pass_threshold is not None:
        passed = overall >= config.pass_threshold

    # A turn slow enough to break the conversation fails the run regardless of
    # how well it scored on everything else.
    failure_reasons = _latency_failures(latencies, config)
    if failure_reasons:
        passed = False

    return NaturalnessResult(
        overall_score=overall,
        overall_label=_label_for_score(overall, config),
        turn_count=len(graded_turns),
        turns=graded_turns,
        conversation_factors=output.conversation_factors,
        factor_averages=_factor_averages(graded_turns),
        label_counts=label_counts,
        summary=output.summary,
        model=target_model,
        pass_threshold=config.pass_threshold,
        passed=passed,
        latency_ms_by_turn=latencies,
        failure_reasons=failure_reasons,
    )