yao/agent/test
Max 4c80c3ffa0 Enhance Script Runner Error Handling and Fail-Fast Logic
- Updated the ScriptRunner to treat both StatusFailed and StatusError as failures, improving failure reporting in test results.
- Modified the fail-fast logic to stop execution on both failure statuses, ensuring quicker feedback during test runs.
- Enhanced error handling in executeTestFunction to recover from panics and provide clearer error messages for JavaScript exceptions, improving test reliability and debugging.
2025-12-22 18:41:13 +08:00
..
assert.go Implement Script Testing Framework and Enhance Test Context Management 2025-12-21 10:44:59 +08:00
context.go Implement Script Testing Framework and Enhance Test Context Management 2025-12-21 10:44:59 +08:00
DESIGN.md Implement Script Testing Framework and Enhance Test Context Management 2025-12-21 10:44:59 +08:00
input.go Enhance Logging Functionality in RequestLogger 2025-12-17 12:33:34 +08:00
interfaces.go Enhance Logging Functionality in RequestLogger 2025-12-17 12:33:34 +08:00
loader.go Enhance Logging Functionality in RequestLogger 2025-12-17 12:33:34 +08:00
output.go Implement Script Testing Framework and Enhance Test Context Management 2025-12-21 10:44:59 +08:00
README.md Implement Script Testing Framework and Enhance Test Context Management 2025-12-21 10:44:59 +08:00
reporter.go Refactor AgentReporter to Use Direct Response Handling 2025-12-18 17:37:37 +08:00
resolver.go Implement Script Testing Framework and Enhance Test Context Management 2025-12-21 10:44:59 +08:00
runner.go Implement Script Testing Framework and Enhance Test Context Management 2025-12-21 10:44:59 +08:00
script.go Enhance Script Runner Error Handling and Fail-Fast Logic 2025-12-22 18:41:13 +08:00
script_assert.go Implement Script Testing Framework and Enhance Test Context Management 2025-12-21 10:44:59 +08:00
script_types.go Implement Script Testing Framework and Enhance Test Context Management 2025-12-21 10:44:59 +08:00
types.go Implement Script Testing Framework and Enhance Test Context Management 2025-12-21 10:44:59 +08:00

Agent Test Framework

A testing framework for Yao AI agents with support for assertions, stability analysis, and CI integration.

Quick Start

Agent Tests

# Test with direct message (auto-detect agent from current directory)
cd assistants/keyword
yao agent test -i "Extract keywords from: AI and machine learning"

# Test with direct message (specify agent explicitly)
yao agent test -i "Hello world" -n workers.system.keyword

# Test with JSONL file (auto-detect agent from path)
yao agent test -i assistants/keyword/tests/inputs.jsonl

# Generate HTML report
yao agent test -i tests/inputs.jsonl -o report.html

# Stability analysis (run each test 5 times)
yao agent test -i tests/inputs.jsonl --runs 5

Script Tests

# Test agent handler scripts (hooks, tools, setup functions)
yao agent test -i scripts.expense.setup -v

# Run specific tests with regex filter
yao agent test -i scripts.expense.setup --run "TestSystemReady" -v

# Run with custom context (authorization, metadata)
yao agent test -i scripts.expense.setup --ctx tests/context.json -v

Input Modes

The -i flag supports three input modes:

1. JSONL File Mode

Load test cases from a file:

yao agent test -i tests/inputs.jsonl

Agent is auto-detected by traversing up from the input file to find package.yao.

2. Direct Message Mode

Test with a single message:

# Auto-detect agent from current working directory
cd assistants/keyword
yao agent test -i "Extract keywords from this text"

# Or specify agent explicitly
yao agent test -i "Hello" -n workers.system.keyword

Output is printed to stdout (or saved to -o if specified).

3. Script Test Mode

Test agent handler scripts (hooks, tools, setup functions):

# Run all tests in a script module
yao agent test -i scripts.expense.setup -v

# Run specific tests with filtering
yao agent test -i scripts.expense.setup --run "TestSystemReady"

# Run with custom context
yao agent test -i scripts.expense.setup --ctx tests/context.json -v

Script test input format: scripts.<assistant>.<module> (e.g., scripts.expense.setupassistants/expense/src/setup_test.ts).

Writing Test Scripts:

Test scripts should be placed alongside the source files with _test.ts or _test.js suffix:

assistants/expense/src/
├── setup.ts           # Source file
├── setup_test.ts      # Test file
├── tools.ts
└── tools_test.ts

Test functions must follow the naming convention Test* and accept (t: testing.T, ctx: agent.Context):

// assistants/expense/src/setup_test.ts
import { SystemReady } from "./setup";

// Test function signature: function Test*(t: testing.T, ctx: agent.Context)
export function TestSystemReady(t: testing.T, ctx: agent.Context) {
  const result = SystemReady(ctx);

  // Use t.assert for assertions
  t.assert.True(result.success, "SystemReady should succeed");
  t.assert.Equal(result.status, "ready", "Status should be ready");
  t.assert.NotNil(result.data, "Data should not be nil");
}

export function TestSystemReadyError(t: testing.T, ctx: agent.Context) {
  // Access context properties
  console.log("Testing with user:", ctx.authorized.user_id);

  // Test error handling
  const result = SystemReady(ctx);
  t.assert.False(result.error, "Should not have error");
}

Available Assertions:

Method Description
t.assert.True(value, msg) Assert value is true
t.assert.False(value, msg) Assert value is false
t.assert.Equal(a, b, msg) Assert a equals b
t.assert.NotEqual(a, b, msg) Assert a not equals b
t.assert.Nil(value, msg) Assert value is null/undefined
t.assert.NotNil(value, msg) Assert value is not nil
t.assert.Contains(s, sub, msg) Assert string contains substr
t.assert.Len(arr, n, msg) Assert array/string length

Test Control:

Method Description
t.Log(msg) Log a message
t.Error(msg) Mark test as failed with msg
t.Fatal(msg) Mark failed and stop test
t.Skip(msg) Skip this test

Command Line Options

Flag Description Default
-i Input: JSONL file path, message, or script ID (required)
-o Output file path output-{timestamp}.jsonl
-n Agent ID (optional, auto-detected) auto-detect
-c Override connector agent default
-u Test user ID test-user
-t Test team ID test-team
--ctx Path to context JSON file for custom authorization -
-r Reporter agent ID built-in
--runs Runs per test (stability analysis) 1
--run Regex pattern to filter which tests to run -
--timeout Timeout per test 5m
--parallel Parallel test cases 1
-v Verbose output false
--fail-fast Stop on first failure false

Agent Resolution

The agent is resolved in the following priority order:

  1. Explicit -n flag: yao agent test -i "msg" -n my.agent
  2. Path-based detection: Traverse up from input file to find package.yao
  3. Current directory: For direct message mode, look for package.yao in cwd

Example directory structure:

assistants/workers/system/keyword/
├── package.yao          <- Agent definition (auto-detected)
├── prompts.yml
├── src/
│   └── index.ts
└── tests/
    └── inputs.jsonl     <- Input file

Input Format (JSONL)

Each line is a JSON object:

{"id": "T001", "input": "Simple text"}
{"id": "T002", "input": {"role": "user", "content": "Message with role"}}
{"id": "T003", "input": [{"role": "user", "content": "Hi"}, {"role": "assistant", "content": "Hello"}, {"role": "user", "content": "Follow-up"}]}
{"id": "T004", "input": "Test", "assert": {"type": "json_path", "path": "field", "value": true}}
{"id": "T005", "input": "Skip this", "skip": true}

Fields

Field Type Required Description
id string Yes Test case ID
input string | Message | []Message Yes Test input
assert Assertion | []Assertion No Assertion rules
expected any No Expected output (exact match)
user string No Override user ID
team string No Override team ID
options Options No Context options (see below)
timeout string No Override timeout (e.g., "30s")
skip bool No Skip this test
metadata map No Additional metadata

Options

The options field allows per-test-case configuration that maps to context.Options:

Field Type Description
connector string Override connector (e.g., "deepseek.v3")
mode string Agent mode (default: "chat")
search bool Enable/disable search mode (default: true)
disable_global_prompts bool Temporarily disable global prompts
metadata map Custom data passed to hooks (e.g., scenario)
skip object Skip configuration (see below)

Options.skip

Field Type Description
history bool Skip history loading
trace bool Skip trace logging
output bool Skip output to client
keyword bool Skip keyword extraction
search bool Skip auto search

Example with options:

{
  "id": "T001",
  "input": "Query users with status active",
  "options": {
    "connector": "deepseek.v3",
    "metadata": {
      "scenario": "filter"
    },
    "skip": {
      "trace": true
    }
  },
  "assert": {
    "type": "json_path",
    "path": "from",
    "value": "users"
  }
}

Using metadata for hook scenarios:

The options.metadata field is passed to agent hooks. For example, a Create Hook can read options.metadata.scenario to select different prompt presets:

{"id": "T001", "input": "...", "options": {"metadata": {"scenario": "aggregation"}}}
{"id": "T002", "input": "...", "options": {"metadata": {"scenario": "join"}}}

Input Types

Type Description Example
string Simple text "Hello world"
Message Single message {"role": "user", "content": "..."}
[]Message Conversation history [{"role": "user", ...}, {"role": "assistant", ...}]

Assertions

Use assert for flexible validation. If assert is defined, it takes precedence over expected.

Assertion Types

Type Description Example
equals Exact match {"type": "equals", "value": {"key": "val"}}
contains Output contains value {"type": "contains", "value": "keyword"}
not_contains Output does not contain value {"type": "not_contains", "value": "error"}
json_path Extract JSON path and compare {"type": "json_path", "path": "$.field", "value": true}
regex Match regex pattern {"type": "regex", "value": "\\d+"}
type Check output type {"type": "type", "value": "object"}
script Run custom assertion script {"type": "script", "script": "scripts.test.Check"}

Assertion Options

Field Type Description
type string Assertion type (required)
value any Expected value or pattern
path string JSON path (for json_path)
script string Script name (for script)
message string Custom failure message
negate bool Invert the result

Examples

JSON path validation:

{
  "id": "T001",
  "input": "What's the weather?",
  "assert": {
    "type": "json_path",
    "path": "need_search",
    "value": true
  }
}

Multiple assertions (all must pass):

{
  "id": "T002",
  "input": "Hello",
  "assert": [
    {
      "type": "json_path",
      "path": "need_search",
      "value": false
    },
    {
      "type": "not_contains",
      "value": "error"
    }
  ]
}

Custom script assertion:

{
  "id": "T003",
  "input": "Test",
  "assert": {
    "type": "script",
    "script": "scripts.test.Validate"
  }
}

Script receives (output, input, expected) and returns:

// Boolean
return true;

// Or detailed result
return { pass: true, message: "Validation passed" };

Negated assertion:

{
  "id": "T004",
  "input": "Hello",
  "assert": {
    "type": "contains",
    "value": "error",
    "negate": true
  }
}

JSON Path Notes

  • Supports dot-notation: $.field.subfield or field.subfield
  • Supports array indexing: field[0], field[0].subfield, field[0].nested[1]
  • Supports multiple expected values (OR logic): "value": ["a", "b"] - passes if actual matches any
  • Auto-extracts JSON from markdown code blocks (```json ... ```)
  • Works with both string output and structured objects

Array index examples:

{"id": "T001", "assert": {"type": "json_path", "path": "wheres[0].like", "value": "%test%"}}
{"id": "T002", "assert": {"type": "json_path", "path": "wheres[0].in[0]", "value": "pending"}}
{"id": "T003", "assert": {"type": "json_path", "path": "joins[0].from", "value": "users"}}
{"id": "T004", "assert": {"type": "json_path", "path": "groups[0]", "value": "category"}}

Multiple expected values (OR logic):

{
  "id": "T005",
  "assert": {
    "type": "json_path",
    "path": "error",
    "value": [
      "missing_schema",
      "missing_query"
    ]
  }
}

This passes if error equals either "missing_schema" or "missing_query".

Output Formats

Determined by -o file extension:

Extension Format Description
.jsonl JSONL Streaming (default)
.json JSON Complete structured
.md Markdown Human-readable
.html HTML Interactive web report

Default Output Path

When -o is not specified in file mode:

{input_directory}/output-{timestamp}.jsonl

Example: tests/output-20241217100000.jsonl

In direct message mode without -o, output is printed to stdout.

Stability Analysis

Run each test multiple times to measure consistency:

yao agent test -i tests/inputs.jsonl --runs 5 -o stability.json

Output includes:

  • Pass rate per test
  • Stability classification (stable, mostly_stable, unstable, highly_unstable)
  • Average/min/max duration
  • Standard deviation

Stability Classification

Pass Rate Classification
100% Stable
80-99% Mostly Stable
50-79% Unstable
< 50% Highly Unstable

Test Environment

The test framework creates a context with configurable environment:

Setting Flag Default
User ID -u test-user
Team ID -t test-team
Locale - en-us
ClientType - test
ClientIP - 127.0.0.1

Priority: Command line flags > Test case fields > Defaults

Custom Reporter Agent

Use -r to specify a custom agent for report generation:

yao agent test -i tests/inputs.jsonl -r report.beautiful -o report.html

The reporter agent receives:

{
  "report": { "summary": {...}, "results": [...] },
  "format": "html",
  "options": { "verbose": true }
}

CI Integration

# Exit code: 0 = all passed, 1 = failures
yao agent test -i tests/inputs.jsonl -o results.jsonl --fail-fast

# Parse JSONL results
cat results.jsonl | jq 'select(.type == "summary")'

# Run script tests
yao agent test -i scripts.expense.setup --fail-fast

GitHub Actions Example

- name: Run Agent Tests
  run: |
    yao agent test -i assistants/keyword/tests/inputs.jsonl \
      -u ci-user -t ci-team \
      --runs 3 \
      -o report.json

- name: Run Script Tests
  run: |
    yao agent test -i scripts.expense.setup -v
    yao agent test -i scripts.expense.tools -v

- name: Run Script Tests with Custom Context
  run: |
    yao agent test -i scripts.expense.setup \
      --ctx tests/context.json \
      --run "TestSystem.*" \
      -v

- name: Check Stability
  run: |
    jq -e '.results | all(.pass_rate >= 80)' report.json

Examples

Agent Tests

# Quick development test (auto-detect agent)
cd assistants/keyword
yao agent test -i "Extract keywords: AI and ML"

# Quick development test (specify agent)
yao agent test -i "Hello" -n workers.system.keyword

# Full test suite with HTML report
yao agent test -i tests/inputs.jsonl -o report.html -v

# Override connector
yao agent test -i tests/inputs.jsonl -c openai.gpt4

# Stability analysis
yao agent test -i tests/inputs.jsonl --runs 10 -o stability.json

# Parallel execution with timeout
yao agent test -i tests/inputs.jsonl --parallel 4 --timeout 2m

# Custom test environment
yao agent test -i tests/inputs.jsonl -u admin -t prod-team

# Custom reporter agent
yao agent test -i tests/inputs.jsonl -r report.beautiful -o custom-report.md

# Full example with all options
yao agent test -i tests/inputs.jsonl \
  -n keyword.agent \
  -c deepseek.v3 \
  -u test-user \
  -t test-team \
  --runs 3 \
  --timeout 10m \
  --parallel 4 \
  -r report.html \
  -o report.html

Script Tests

# Run all tests in a script module
yao agent test -i scripts.expense.setup -v

# Run specific tests with regex filter
yao agent test -i scripts.expense.setup --run "TestSystemReady"

# Run tests matching a pattern
yao agent test -i scripts.expense.setup --run "TestSystem.*" -v

# Run with custom context (authorization, metadata, etc.)
yao agent test -i scripts.expense.setup --ctx tests/context.json -v

# Run with specific user/team
yao agent test -i scripts.expense.setup -u admin -t ops-team -v

# Combine options
yao agent test -i scripts.expense.setup \
  --ctx tests/context.json \
  --run "TestSystem.*" \
  --timeout 30s \
  -v

Exit Codes

Code Description
0 All tests passed
1 Tests failed, configuration error, or runtime error