Test privately
Create a fixed test version, run free private tests on real inputs, and read the report, tool trail and evidence.
Tests run one fixed version of your agent on real inputs and live chain data, exactly as a buyer's run would, but privately and free. This page covers creating a test version, running tests, their limits, and reading the result and its evidence. It is for creators checking an agent before review or publication.
Create a test version#
A test always runs a fixed version, never your editable draft.
- In the workflow builder, choose Save & create test version.
- On a template, choose Save & test.
- Importing an SDK package creates the first test version for you.
"Creating a test version fixes this exact workflow and its hash. Later edits never change it." Each version of an agent is numbered (version 1, 2, 3 and so on) and shows the first characters of its hash.
Before the version exists, the server checks that every tool is approved and covered on every declared chain, and that any AI analysis or sandboxed step can run here. If not, nothing is created and you see "This workflow cannot become a test version:" followed by the reasons (for template agents, "This agent cannot become a test version:").
Open the test page#
The test page opens after you create a version. You can also reach it from:
- the Test stage, which shows "Choose a workflow version to test" when no version is selected;
- Test version N in the Versions, tests and run outcomes table on Versions.
Workflow versions are tested at /studio/workflow/test; template agents at /studio/test. Watchtower versions are tested by creating a monitor.
Run a test#
- Under New test run, choose the Analysis chain (one of the version's chains).
- Fill in the inputs. Each field shows the input's description and type; optional inputs are marked "(optional)".
- Choose Run test. "Free, bounded creator test. It uses the version's call budget and records every read as evidence."
- The run appears under Test runs (inputs, chain and start time in UTC) with its status, and its result opens under Run result when it finishes.
Tests are free, private and never listed. They never earn revenue and never count as marketplace runs.
| Limit | Value | Message |
|---|---|---|
| Runs in progress | 3 per account at a time | "The development queue is full. Wait for a run to finish." |
| Runs per day | 20 per account in a rolling 24 hours, counting tests and paid runs you buy | "The bounded daily development budget is exhausted. Try after the rolling 24-hour window resets." |
| Version status | A paused, revoked or deprecated version accepts no tests | "This version does not accept new tests." |
A platform-wide queue limit also applies; if it is full, wait and try again.
Read the result#
The result shows the run id, its status and how many RPC calls it used, then the report.
The report header#
"Report complete", "Report partial" or "Report failed", then a usage line such as "9/16 RPC calls · 2 tool executions · 1840 ms · provider cost unpriced", and the version number and hash the report is bound to.
Report sections#
Each declared section shows its status (complete, partial, skipped or failed), a "required" tag if it is required, and its findings. A finding shows its value, or its status and reason if it has no value. Values from inputs are marked "(from input, not observed)". Under each finding, evidence chips link to the evidence items it cites. See Report sections and findings.
Tool and evidence trail#
A table of every step in run order:
| Column | Shows |
|---|---|
| Step | The step's id. |
| Type | input, tool, condition, transform, report-section, ai-analysis, sandboxed-transform or output. |
| Tool | The tool and version a data step ran. |
| Status | succeeded, skipped, failed or disabled; for conditions also "condition was true", "false" or "unknown". |
| Calls | Reads the step made. |
| Evidence | Chips for the evidence the step captured. |
| Detail | Why a step was skipped or failed, for example "Dependency metadata was failed." |
Evidence#
"Evidence (N)" lists every item the run captured. Each line shows the evidence id, tool, method, completeness (complete-for-request, truncated, empty-result, unavailable or error) and the block it was pinned to. Expand an item to see its provider and trust level, chain, confirmation, the time it was observed, its content hash, and the exact request and result. See Evidence.
Errors and limitations#
Errors are listed as "step: category: message", for example a permission refusal or a read-limit stop. The report's limitations follow. Fix the cause in the draft and create a new version.
When a test counts#
A test counts for review, publishing and saved examples when the run succeeded and its report is complete and meets the workflow result contract (every required section complete and backed by stored evidence). For Token Researcher, a partial test that verified the token's metadata can also be used, with your explicit consent in review.
On a workflow's test page:
- Pricing links to Set its price and publish it.
- Submit for review pins one succeeded test to the version: choose it under Creator test to submit and choose Submit for review. The version moves to In review, and the page confirms "Pinned to creator test (id) and version hash (hash)." Continue with Continue to platform checks and operator review. A version can pin only one test; to submit different evidence, create a new version. See Review and the badge.
Test well#
- Test every chain you declare, with inputs a buyer is likely to use.
- Test edge cases: an address with no contract, a token without an owner, a quiet pool. Check that the sections you marked required are still complete.
- Check the RPC calls of each test. A buyer's quote covers at most 16 metered reads; a version that needs more cannot complete a paid run.
- Open the evidence behind each finding and make sure it proves what the finding says.
Template agent tests#
Template agents use /studio/test: pick the Fixed version, choose the Analysis chain, fill in the kind's inputs (for example Token contract address, or Sender address, Target address, Calldata and Native value · atomic units for the Transaction Inspector), and choose Run private test. Paste unsigned payloads only: "Never paste a private key or signed transaction." Recent test runs lists your tests, and Continue to review opens review for the version.