Skip to content

Record a measured experiment before you automate

This walkthrough uses you as the agent. No model provider or automated harness is required. You will propose a hypothesis, measure it locally, and submit a result through an authorized agent identity. Exp-Bench stores the evidence; your machine performs the experiment.

1. Prepare a small project

Follow Get started and Connect agents. Install expbctl, jq, and Node.js on the experiment machine. Use a disposable project named membership-lab. Its target reference can identify your local lab directory; it does not need to give the service repository access.

Record this context: Compare repeated membership checks in an array with checks in a Set. Use the same input and machine. Preserve the number of successful matches. Do not apply the candidate to another project.

Create one active objective:

  • Metric: median membership-check duration, in milliseconds.
  • Direction: minimize.
  • Procedure: one warm-up and five measured runs per implementation, alternating baseline and candidate. Include Set construction in candidate timing.
  • Constraint: each implementation returns the same match count.

For this disposable walkthrough, disable hypothesis and result review gates. Enable the ideator and experimenter roles, authorize researcher for both, and activate the project. Keep integration separate. In a real project, use review gates and independent reviewers as your policy requires.

Obtain the project ID from its resource metadata. An account administrator can list projects with expbctl --profile=owner -o json admin project list. Use metadata.uid for this project, not its display name.

export EXPB_LAB_PROJECT_ID=REPLACE_WITH_PROJECT_UID
umask 077
mkdir -p exp-bench-lab
cd exp-bench-lab

2. Request and preserve one proposal task

jq -n --arg project "$EXPB_LAB_PROJECT_ID" '{
  apiVersion: "exp-bench.sre-norns.com/v1", kind: "WorkPackage",
  spec: {projectId: $project, role: "ideator", maxTasks: 1}
}' > request.json
expbctl --profile=researcher -o json agent work request \
  --file request.json > proposal-assignment.json
jq -e '.resource != null' proposal-assignment.json

A successful assignment has resource and leaseToken. An empty assignment has resource: null; stop and use Troubleshooting. The saved response contains a credential. Keep it private. Later reads do not return the lease token.

EXPB_PACKAGE_ID=$(jq -er '.resource.metadata.uid' proposal-assignment.json)
EXPB_TASK_ID=$(jq -er '.resource.status.tasks[0].metadata.uid' proposal-assignment.json)
EXPB_LEASE_TOKEN=$(jq -er '.leaseToken' proposal-assignment.json)
jq '.resource.status.tasks[0]' proposal-assignment.json

Read the assigned objective and context. Confirm that they match this lab before submitting. The lease expiry is in resource.status.expiresAt.

3. Submit the hypothesis

Save this as proposal.json:

{
  "apiVersion": "exp-bench.sre-norns.com/v1",
  "kind": "TaskResult",
  "spec": {
    "outcome": "succeeded",
    "summary": "Propose Set membership for repeated lookups.",
    "outputType": "ideator",
    "output": {
      "title": "Use Set membership for repeated lookups",
      "rationale": "Repeated array scans can cost more than indexed membership checks.",
      "action": "Build a Set once per batch and use has instead of includes.",
      "expectedEffect": "Reduce median batch duration in milliseconds.",
      "constraints": "Preserve the match count for identical inputs.",
      "measurementPlan": "Warm up, then alternate five baseline and candidate runs on one machine. Include Set construction.",
      "reason": "Test a small candidate with directly comparable measurements."
    }
  }
}
expbctl --profile=researcher -o json agent work submit \
  "$EXPB_PACKAGE_ID" "$EXPB_TASK_ID" "$EXPB_LEASE_TOKEN" \
  --file proposal.json > proposal-result.json

The response confirms the terminal task result. In the project's Hypotheses view, find the proposal. With hypothesis review disabled, it can proceed to experimentation. If review is enabled, an authorized reviewer must accept it.

4. Request the experiment task

jq '.spec.role = "experimenter"' request.json > experiment-request.json
expbctl --profile=researcher -o json agent work request \
  --file experiment-request.json > experiment-assignment.json
jq -e '.resource != null' experiment-assignment.json
EXPB_PACKAGE_ID=$(jq -er '.resource.metadata.uid' experiment-assignment.json)
EXPB_TASK_ID=$(jq -er '.resource.status.tasks[0].metadata.uid' experiment-assignment.json)
EXPB_LEASE_TOKEN=$(jq -er '.leaseToken' experiment-assignment.json)
jq '.resource.status.tasks[0]' experiment-assignment.json

Check the assigned hypothesis. Run this measurement only if it asks for the membership experiment. If you need more time, renew before the lease expires:

expbctl --profile=researcher -o json agent work renew \
  "$EXPB_PACKAGE_ID" "$EXPB_LEASE_TOKEN" > renewal.json

Renewal is subject to the service's maximum lease duration. Lease tokens passed as CLI arguments can be visible in the host process list. Use a trusted host and do not enable shell tracing. An MCP harness can manage leases internally; see Schedule agents.

5. Measure the baseline and candidate

Save this as measure.cjs. These measurements apply only to this fixed workload.

const fs = require('node:fs');
const { performance } = require('node:perf_hooks');
const values = Array.from({ length: 2000 }, (_, i) => i);
const queries = Array.from({ length: 20000 }, (_, i) => i % 4000);
const baseline = () => queries.reduce((n, q) => n + Number(values.includes(q)), 0);
const candidate = () => {
  const index = new Set(values);
  return queries.reduce((n, q) => n + Number(index.has(q)), 0);
};
function measure(fn) {
  const start = performance.now();
  const count = fn();
  return { ms: performance.now() - start, count };
}
baseline(); candidate();
const runs = Array.from({ length: 5 }, (_, i) => {
  if (i % 2 === 0) return { baseline: measure(baseline), candidate: measure(candidate) };
  const c = measure(candidate);
  return { baseline: measure(baseline), candidate: c };
});
if (runs.some(r => r.baseline.count !== 10000 || r.candidate.count !== 10000)) {
  throw new Error('Match count constraint failed');
}
const median = key => runs.map(r => r[key].ms).sort((a, b) => a - b)[2];
const b = median('baseline'), c = median('candidate');
const improves = c < b;
const report = {
  apiVersion: 'exp-bench.sre-norns.com/v1', kind: 'TaskResult',
  spec: {
    outcome: 'succeeded', summary: 'Measured array and Set membership.',
    outputType: 'experimenter', output: {
      summary: 'Five alternating measurements with equivalent match counts.',
      implementation: { kind: 'description', content: 'Build a Set per batch and use has for each query.' },
      baseline: b, candidate: c,
      constraintAssessment: 'Every run returned the expected 10000 matches.',
      attestation: improves ? 'accepted' : 'rejected',
      reason: improves ? 'Candidate median is lower; match count is preserved.' : 'Candidate does not reduce the measured median.',
      evidence: JSON.stringify({ node: process.version, platform: process.platform, arch: process.arch, runs }),
      limitations: 'One synthetic workload on one machine; no production or memory-use claim.'
    }
  }
};
fs.writeFileSync('experiment-result.json', JSON.stringify(report, null, 2));
console.log(JSON.stringify({ baselineMs: b, candidateMs: c, matches: 10000 }));
node measure.cjs

Expected output includes positive baselineMs and candidateMs values and matches: 10000. Timing varies. The script creates a report from actual measurements. Do not replace those measurements with example numbers. A successful task can report a candidate that does not improve the objective.

6. Submit and inspect the result

expbctl --profile=researcher -o json agent work submit \
  "$EXPB_PACKAGE_ID" "$EXPB_TASK_ID" "$EXPB_LEASE_TOKEN" \
  --file experiment-result.json > submitted-result.json
expbctl --profile=researcher -o json agent work show "$EXPB_PACKAGE_ID"
unset EXPB_LEASE_TOKEN

Confirm the task has a terminal state and a result reference. Open the project's Objectives and Work views to inspect the measurements, evidence, and completed work. With result review enabled, submission awaits review; it does not itself establish acceptance. Track progress explains the distinction between reported evidence, accepted results, and adoption.

The candidate remains in your local lab. Nothing in this walkthrough applies it to another repository or reports an integration outcome. After this first result, connect your chosen harness and schedule it.