Development
Build and check the project with Cargo:
cargo build
cargo testRun the local binary without installing:
cargo run --
cargo run -- run "summarize this repository"Use the local binary to test the interactive TUI, automation mode, configuration, and tools behavior while developing.
RHO_FIRST_RUN opens the setup screen on a machine that already has config. It does not clear history or credentials.
RHO_FIRST_RUN=signin rho
RHO_FIRST_RUN=model rho
RHO_FIRST_RUN=1 rhosignin opens the provider menu. model opens the model list. 1 opens whichever step a real first launch would open. On a configured machine that already lists models, 1 goes to the model step. To see the signed-out session, run /logout for the active provider.
Set RHO_LOG to a tracing env-filter to print spans on stderr. It is off by default. Examples: RHO_LOG=rho=info or RHO_LOG=rho=debug.
Local validation
Use the fast workflow while editing. It always checks formatting and architecture, then checks one package without compiling every target:
python3 scripts/validate.py fast --package rho-sdkYou can add a library test, an integration test, or a test-name filter:
python3 scripts/validate.py fast --package rho-sdk --lib
python3 scripts/validate.py fast --package rho-coding-agent --test automation_cli
python3 scripts/validate.py fast --package rho-coding-agent --test automation_cli --filter streams_json_eventsRun the full workflow before opening or updating a pull request:
python3 scripts/validate.py fullThe full mode runs policy and script checks, Clippy for all targets and features, normal workspace tests, documentation tests, SDK feature and downstream checks, and the docs TUI proof-plate PTY check. It stops at the first failure. Both modes cap Cargo at 8 jobs. A lower CARGO_BUILD_JOBS value remains in effect. They also set RUST_MIN_STACK to at least 16 MiB so rustc 1.92 does not overflow its 8 MiB compile-thread stack.
Development and test profiles use reduced debug information to keep artifacts and link times smaller while retaining line-number backtraces. Set CARGO_PROFILE_DEV_DEBUG=2 or CARGO_PROFILE_TEST_DEBUG=2 when a debugging session needs full symbols.
Crate publish preparation
CI runs scripts/check_crate_publish_prep.sh before release packaging. That script calls scripts/crate_publish_prep.py, which dry-runs cargo publish for each independently released crate.
Internal path dependencies use this registry-boundary rule:
- If every internal dependency's exact workspace version already exists on crates.io (including yanked releases), verify against the registry. Do not path-patch any dependency.
- If any direct internal dependency version is not on crates.io yet, path-patch the full transitive internal dependency closure of the package under test. Path-source crates keep
path =edges to siblings; patching only the unpublished leaf would load a second copy of shared crates such asrho-sdkand break type identity. The closure patch keeps one graph so a coordinated same-PR version cut can still package before dependencies are published. - crates.io transport or HTTP failures fail the check. A timeout or 500 is not treated as "unpublished".
This catches the failure mode where a refactor moves public symbols into a dependency without cutting that dependency's version: local workspace builds still pass, but cargo publish verification against crates.io does not.
scripts/publish_workspace_crates.sh --dry-run uses the same patch selector. Real publishes never path-patch; they wait for each dependency version to index on crates.io first.
Local checks:
python3 scripts/crate_publish_prep.py --self-test
./scripts/check_crate_publish_prep.shThe self-test uses fixtures/publish-boundary, where a consumer imports a symbol that exists only on the workspace copy of a dependency. Without a workspace path patch, verification fails the same way a real crates.io publish would.
Thermo-nuclear review workflow
The repository includes .rho/workflows/thermo-nuclear-review/ for a deep review of the current branch. It collects one bounded Git context pack, runs three read-only review lanes in parallel, then sends their structured findings to one worker that applies safe fixes. If the selected scope has no changes, the workflow takes a no-op path instead of starting review agents.
rho workflow validate .rho/workflows/thermo-nuclear-review/workflow.star
rho workflow plan .rho/workflows/thermo-nuclear-review/workflow.star \
--input 'base="main"' \
--input 'scope="all"'base must resolve exactly; an invalid explicit ref does not fall back to a different branch. scope accepts all, committed, or uncommitted. See .rho/workflows/thermo-nuclear-review/README.md for its graph, inputs, run commands, and context collector test.
Test selection
Before adding, expanding, reviewing, or deleting tests, follow the project skill rho-test-selection at .agents/skills/rho-test-selection/SKILL.md. It defines the failure-mode / owner-layer gate, Tier A/B/C rules, determinism requirements, and PTY-as-product-gate defaults.
Pull requests that add tests should fill the test-gate section in the pull request template.
Test prerequisites
Some integration tests spawn local fixtures and need host tools on PATH:
stdio_lifecycle_and_failure_isolationincrates/rho/src/tools/mcp_tests.rsis Unix-gated and requirespython3to run its stdio MCP server fixture.
Interactive TUI PTY harness
Rho includes a deterministic PTY harness in crates/rho-tui-pty for automated interactive TUI tests. Prefer it over manual Herdr smoke tests for regressions that can be expressed as scripted scenarios. PTY is the product gate for interactive behavior; unit tests under crates/rho/src/tui stay limited to pure logic.
Layers
- PTY controller - spawn a selected
rhobinary in a pseudo-terminal, inject keys/paste/mouse, resize, drain output, and kill-on-drop - Screen model - reconstruct the visible terminal with a VT parser and assert user-visible text
- Scenarios - named action/assertion sequences over
RHO_TUI_TEST_MODE=matrix - Artifacts - on failure, keep raw PTY bytes, reconstructed screen, action log, and redacted env
The controller supports Unix PTYs and Windows ConPTY. The broader scenario suite still targets Unix; a shared native smoke test exercises startup, multiline Unicode paste, composer navigation, submission, and exit on Windows as well. The harness compiles crates/rho/src/pty.rs by path instead of depending on rho-coding-agent, so the Cargo workspace (and release-please's cargo-workspace plugin) stay acyclic.
Run harness self-tests
cargo test -p rho-tui-ptyRun the CI smoke scenarios
cargo test -p rho-coding-agent --test tui_ptySmoke scenarios cover startup/stream/exit, cancel-and-resubmit, resize-during-stream, scroll-during-stream, and terminal restoration.
On Linux, the startup smoke test launches Rho with a 1 MiB main-thread stack, matching the Windows linker default. This catches debug async-frame overflows that Linux's larger default stack would hide.
The native smoke test runs in the Linux, macOS, and Windows workspace CI jobs:
cargo test -p rho-coding-agent --test tui_pty_native --jobs 8Windows uses ConPTY, which requires Windows 10 version 1809 or later. This focused test does not replace the Unix scenario suite or prove every Windows terminal interaction.
Run one scenario locally
cargo build -p rho-coding-agent
cargo run -p rho-tui-pty --bin rho-pty-scenario -- --list
cargo run -p rho-tui-pty --bin rho-pty-scenario -- --bin target/debug/rho startup_stream_exit
cargo run -p rho-tui-pty --bin rho-pty-scenario -- --bin target/debug/rho --smoke
cargo run -p rho-tui-pty --bin rho-pty-scenario -- --bin target/debug/rho --timing startup_stream_exit
cargo run -p rho-tui-pty --bin rho-pty-scenario -- --bin target/debug/rho --timing startup_first_frame--timing records wait latency samples. For first-frame paint, compare the named wait_for_text:rho sample rather than mixed p50/p95/p99 across unrelated waits. Do not treat those milliseconds as a CI budget.
Failure artifacts default to a temp directory (or --artifacts <dir>). Successful runs do not retain artifacts.
Regenerate the docs TUI proof plate
Dark and light proof plates are captured from one matrix-mode PTY session (not hand-drawn), then rendered with two SVG palettes:
- dark (README + site):
docs/assets/rho-ui-demo.svganddocs/public/assets/rho-ui-demo.svg - light (site only):
docs/public/assets/rho-ui-demo-light.svg
After TUI layout or chrome changes, regenerate those paths:
bash scripts/check_docs_ui_demo.sh --writeDetect drift without writing:
bash scripts/check_docs_ui_demo.sh --checkThe CI quality job and python3 scripts/validate.py full run the check. The generator needs a Unix PTY and a debug build because RHO_TUI_TEST_MODE=matrix is debug-only. Load-volatile fragments (tool durations and statusline usage) are pinned in the SVG. The header package version is pinned at capture input via RHO_TUI_DISPLAY_VERSION so release bumps do not force a regen.
Environment isolation
Scenarios launch Rho with:
- temporary
HOMEand--config RHO_TUI_TEST_MODE=matrix(debug builds only)- host terminal markers stripped (
TMUX,TERM_PROGRAM, Herdr vars, editor markers, and related identity env) check_for_updates = falseandweb_search_provider = "disabled"in the isolated config
When to use Herdr instead
Use the Herdr sibling-pane workflow for exploratory checks, novel bugs that are not yet encoded as scenarios, or parity checks against a real terminal renderer. See the Herdr page and the rho-tui-pty-testing and rho-tui-herdr-testing skills.
Workflow limit receipts
Planning budgets come from a checked-in receipt, not from a guessed constant. The planner reads that receipt. A deterministic generator builds separate stress cases for source modules, evaluator work and heap, values, a 750-node and 7,500-edge graph, schemas and conditions, serialized graph size, runtime output, templates, prompts, argv, inputs, and planner process frames.
The receipt and corpus map are crates/rho/src/workflow/fixtures/limit_receipt.json and crates/rho/src/workflow/fixtures/limit_corpus.json. The generator is scripts/workflow_limit_corpus.py.
Verify after a build:
cargo build -p rho-coding-agent -j 8
python3 scripts/measure_workflow_limits.py --rho target/debug/rhoThe command runs each generated case in the product planner worker, reads the worker's evaluator tick and peak-heap counters, derives program and runtime values from the returned plan, and also runs workflow validate. Both paths isolate HOME and RHO_HOME so local agents and credentials cannot affect the corpus. It fails if a deterministic value differs from the receipt, if a process frame differs, or if wall time or address space loses its stated safety margin. Wall time and address space use checked baselines because OS load can change them. The verifier allows at most twice the baseline and still requires the separate minimum margin in the receipt. Add --json-output /tmp/workflow-measurements.json to retain completed measurements even when comparison fails.
The evaluator heap metric matches Starlark's enforced quantity: mutable-heap peak plus frozen-heap allocated bytes. Remeasuring the unchanged corpus found 67,111,744 mutable bytes and 19,176 frozen bytes, totaling 67,130,920 bytes. The old 50,334,528-byte receipt omitted frozen allocations and understated the current mutable peak by 16,777,216 bytes. The accepted heap budget is now 80,796,392 bytes, retaining the existing 13,665,472-byte absolute margin. No corpus workload was reduced. graph_bytes measures the entire serialized scoped program, including its root and input schemas, rather than just the root graph.
On Linux, the address-space value is the highest /proc/<pid>/status VmSize seen after the supervised child starts the planner worker. That omits the short period before the child applies its limit. The checked debug build used 1,170,087,936 bytes under a 4,294,967,296-byte OS ceiling. That ceiling is a coarse process backstop, RLIMIT_AS on Linux and the same value as a Job Object process-memory commit limit on Windows. It is not the product tripwire. Product memory policy lives in the receipt. Virtual size is much larger than resident memory because allocators reserve address space without committing it. If the worker needs more than the checked amount, the check reports the measured value and the hard limit.
environment_expansion_bytes is a schema sentinel, not a corpus measurement. Workflow schema v1 forbids source-controlled environment entries and keeps a one-byte accepted floor.
Current measured stress values include 750,000 source bytes, 75 modules at depth 15, 750,019 evaluator ticks, 67,130,920 evaluator heap bytes, 750 nodes, 7,500 edges, a 756,418-byte schema, a 7,515,385-byte program, 6,291,456 bytes per retained stream, and 50,331,648 total retained bytes. Read the receipt for every value and margin.
Cancellation uses a separate measurement. It starts a real workflow owner, waits on a Unix socket until a compiled command node is active, and starts a second Rho process to run workflow cancel. Linux pidfd_open checks that the command process has exited. Process completion, not a sleep, ends each wait.
python3 scripts/measure_workflow_cancellation.py \
--rho target/debug/rho --repeat 5That command needs Linux with Unix sockets and pidfd_open, and rustc on PATH. It uses a new temporary RHO_HOME for each sample. The checked run measured 33 ms for acknowledgement, final command cleanup, and workflow owner completion. The accepted limits are 2,000 ms, 2,000 ms, and 2,500 ms. The cancellation command checks both the accepted limits and twice the checked baseline.
Compaction replay eval
Live compaction metrics show how often compaction runs and what it costs. They cannot show what a summary lost. The offline replay eval answers that. Use it before you change summarizer prompts or models, or compaction defaults. It is a local tool: it runs on your own saved sessions and never in CI.
For each replay point, the eval runs the production compactor on a saved session's uncompacted history. It then asks fixed probe questions from the compacted context and scores each answer against the full transcript. References come from the whole history, not just the removed span, so every variant answers the same questions at the same point:
| Probe | Reference | Scoring |
|---|---|---|
files_changed | paths named by write and edit-tool calls | exact match on path suffix |
test_result | the latest test, build, or check command and its status | judge model |
user_requests | the latest user messages | judge model |
errors | the latest failed tool calls | judge model |
A probe is skipped when the history has nothing for it. files_changed sees only write and edit-tool calls. Files changed by shell commands, such as a mv or a generated file, are not in its reference, so it measures part of what changed. Naming extra paths costs nothing. Run a --tiers none variant alongside the others. It answers from the uncompacted history, so its score is the ceiling that answer and judge noise allow. The judge sees only the question, the reference facts, and the answer, never the whole transcript. Pin the answer and judge models within a comparison; changing either one changes the scores. The report records the probe-set version, so only compare reports that share it.
cargo build -p rho-coding-agent -j 8
python3 scripts/compaction_eval.py --recent 8 --points 2 \
--rho-args '--provider openai-codex --model gpt-6-luna --auth codex' \
--variant 'none=--tiers none' \
--variant 'default=--tiers text --summarizer session' \
--variant 'target30=--tiers text --summarizer session --target-percent 30'Each --variant passes arguments to the hidden rho __compaction_eval command. You can set threshold and target percents, a fixed --context-window, a --summarizer provider/model (or session), --tiers text to turn off native compaction, --tiers none for the ceiling, and --answer-model and --judge-model. Root --provider and --model choose the session model. Without --context-window, each point's window is sized so the point sits exactly at the threshold, as a live automatic compaction would. Points below 32,768 estimated tokens are skipped. Across 221 local sessions, the median peak before any compaction was about 51,000 tokens, and 131 sessions reached 32,768. Replay stops at a session's first compaction, so an earlier summary never counts toward the score.
The script prints a Markdown table: mean probe scores, post-compaction size as a fraction of the original, summary output tokens, cost, and latency. Each point in a report also carries the text summary the compaction wrote (summary, empty for native compaction and for points that needed no summary), so you can read what a configuration kept. Reports are saved under --out (default /tmp/rho-compaction-eval) with private permissions. --render-only redraws the table from saved reports.
Saved sessions contain your code and any secrets you pasted or printed. The eval sends them to the models you configure, the same way the original session did. Reports quote transcript excerpts, so keep them local and never commit them. Eval requests do not reach the usage ledger or compaction_events, and elided tool results go to a temporary directory that is deleted after each point.
Provider identity and auth modes
A provider identifies one API or product surface. If two login methods use the same API base, wire protocol, and model catalog, add both to that provider's auth_modes list rather than adding a second provider. The first mode is the default. Keep separate providers when endpoints, protocols, catalogs, or product surfaces differ. For example, OpenRouter API-key and OAuth access share openrouter, while the OpenAI API and Codex remain openai and openai-codex.
When retiring a same-API provider name, add a load-time alias that selects both the canonical provider and the matching auth mode. New config, model references, favorites, runtime identities, and cache entries must use the canonical provider name.
Model integration layers
Model integrations are split into three layers:
crates/rho-providers/src/model/defines provider registry, catalog, and application model support without wire types.crates/rho-providers/src/protocol/converts the canonical SDK model to and from API wire formats. OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and Google Gemini Generate Content are implemented here.crates/rho-providers/src/providers/owns credentials, endpoint selection, headers, retries, continuation state, and transport policy for each provider. Multiple providers may consume one protocol codec.
Keep provider-specific fields in protocol or provider modules unless the agent needs the underlying concept. Adding a protocol stub does not make a provider available: provider identity, authentication, model discovery, runtime construction, and documentation must be implemented separately.
Architecture guardrails
Run the lightweight architecture checks before submitting structural changes:
python3 scripts/check_architecture.py
python3 scripts/check_architecture.py --self-testThe script uses only the Python standard library and reads policy from scripts/architecture.json. It enforces these repository policies:
- Hand-written production Rust files under workspace crate
src/directories, plus cratebuild.rsfiles, have a 1,000-line default budget (default_production_line_budget). - Dedicated test files, including files under a
tests/directory and*_test.rs,*_tests.rs, ortests.rs, are excluded. Inline tests still count toward their production file's budget. - Generated Rust files are excluded only when their exact path and reason are recorded in
generated_filesinscripts/architecture.json. There are currently no generated-file exclusions. - Existing oversized production files are listed explicitly in
legacy_file_budgets. Their ceilings prevent further growth and should be lowered or removed as the files are split. crates/rho-providers/src/credentialsmust remain independent ofmodel, keeping credential storage separate from model runtime metadata (forbidden_dependencies).crates/rho/src/main.rshas a 50-line thin-binary budget so application orchestration remains in the library (thin_binary_budgets).- Package dependency boundaries are also declared in
forbidden_package_dependenciesso lower-level crates cannot depend on the application or reverse the intended layering.
Current legacy file budgets are: none.
Do not raise a budget just to make a check pass. Prefer extracting a cohesive module and reducing the recorded ceiling. If a generated file must be added, list its exact repository-relative path with a non-empty reason so the exclusion remains reviewable. When changing the scanner or policy, update its self-tests and this documentation together.
Rust toolchain and MSRV
The rho-sdk minimum supported Rust version (MSRV) is 1.86. The rho-coding-agent application MSRV is 1.92 because its terminal, credential, and terminal-native Mermaid rendering dependencies require a newer compiler. Both values are declared as package.rust-version in Cargo metadata and tested in CI.
rustc 1.92.0's default compilation-thread stack is 8 MiB. Type-checking this workspace's dependency graph (and some individual crates such as hashbrown, syn, toml_edit, and num-traits) overflows that stack. rustc reports it as SIGSEGV plus RUST_MIN_STACK=16777216, or as a query ICE (active query job entry / DepNodeIndex pack assert). .cargo/config.toml sets RUST_MIN_STACK to 16 MiB when the variable is unset. validate.py raises any smaller value to that floor and caps Cargo at 8 jobs. Override the job cap with CARGO_BUILD_JOBS or cargo -j.
When either MSRV changes, update the matching Cargo manifest, this section, and CI together. An MSRV increase must not ship as a patch release. On a stable major line, an SDK MSRV increase requires at least a minor version increase and must be called out in release notes. Emergency compiler requirements caused by a security or soundness fix may skip normal notice, but still require coordinated metadata and CI updates.
Embedders only need the SDK crate's declared package.rust-version. See SDK installation.