Skip to content

The API coordinates jobs and Workers execute them

Urth separates the control plane from probe execution. The API server owns resource management, authorization, placement, and execution records. Workers run on customer-managed infrastructure and initiate their connections to the API and broker.

Components and connections

Scroll wide diagrams horizontally to read all of their labels.

flowchart TB
    Client["Web interface / urthctl"] -->|Resource API| API["API control server"]
    API <-->|Store resources| DB["PostgreSQL"]
    API -->|Relay committed jobs| Queue["NATS JetStream: Runner queues"]

The control plane runs at the configured https://urth.example.org address. PostgreSQL stores the Result and dispatch records. The relay publishes committed jobs to NATS. The customer manages the Worker and its target network access:

flowchart TB
    Worker["Customer-managed Worker"] -->|Outbound HTTPS| API["API: enroll, claim, report"]
    Worker -->|Outbound broker connection| Queue["NATS: pull and acknowledge"]
    Worker -->|Probe traffic| Target["Target service"]

The second diagram's arrows describe connections initiated by the Worker. The API and broker return jobs and responses over those connections.

Component Role
API control server Authenticate callers, manage resources, select Runners, authorize claims, and accept reports.
PostgreSQL Store resources, Results, and transactional dispatch records.
NATS JetStream Hold job messages in logical Runner queues for Workers to collect.
Runner Define the stable placement, admission, and authorization boundary for a queue.
Worker Collect and claim jobs, execute probes, and upload results and artifacts.
Web interface and urthctl Manage and inspect the same API resources.

A Runner is a logical resource, not a server that Urth starts. A WorkerInstance records an enrolled Worker and its session. The Worker process remains the customer's responsibility.

Follow a job through the system

%%{init: {"sequence": {"useMaxWidth": false, "actorMargin": 20, "width": 100, "wrap": true}}}%%
sequenceDiagram
    participant U as User / CLI
    participant A as API
    participant Q as Runner queue
    participant W as Worker
    participant S as Service
    U->>A: Trigger Scenario
    A->>A: Select Runner and commit dispatch
    A->>Q: Relay committed dispatch
    W->>Q: Pull job
    Q-->>W: Dispatch identifiers
    W->>A: Authenticated claim
    A-->>W: Execution input and run capability
    W->>Q: Acknowledge claim
    W->>S: Execute probe
    S-->>W: Service response
    W->>A: Report Result and Artifacts
    U->>A: Inspect Result

The Worker pulls queue messages and calls the API itself. The service response is the measurement; it is separate from Worker liveness.

  1. A user triggers a Scenario through the web interface, CLI, or API.
  2. The control plane selects an active, authorized Runner that meets placement requirements.
  3. It records the Result and dispatch in one PostgreSQL transaction.
  4. A relay publishes the committed dispatch to the Runner's JetStream queue.
  5. A Worker pulls the message and requests an authenticated claim from the API.
  6. The API checks current Runner, Worker, and project authorization before returning execution input.
  7. The Worker acknowledges the dispatch, executes the probe, and reports the Result and Artifacts.

Queue messages contain dispatch identifiers, not the executable script or probe credentials. Execution input arrives after an authorized claim. The Worker receives bounded authority to report that execution.

The dispatch record allows a broker outage to delay publication without losing the committed job. It does not guarantee successful or exactly-once execution. A Worker can fail during a probe. Inspect Result and dispatch diagnostics when reporting or execution fails.

Understand enrollment and job authority

A Runner machine token authorizes enrollment. The Worker proves possession of its persistent installation key and receives a Worker session and restricted broker credentials. It uses the session to claim work. A run capability permits status and Artifact writes for the claimed execution within its lifetime.

These credentials serve different purposes. A project grant is checked at placement and again at claim. Matching labels do not bypass authorization. This guide describes the current source model; production security validation remains a release gate.

One queued job has one executor

Workers sharing a Runner compete for its jobs. Each Result is placed on one Runner and claimed by one Worker. Adding more Workers to the same channel provides execution capacity and resilience to process loss. It does not make every Worker run every probe.

For distinct measurement locations, use separate Runners and explicitly targeted Scenario runs. See Probe from several locations.