# Observatory Agent Guide

Observatory is a multi-user cost and operational-metrics app. Each
signed-in user has an isolated set of provider connectors. The dashboard is
read-only; a user's agent configures connectors through the authenticated API.

## API

- `GET /api/connectors` — connector catalog and the current user's connection status.
- `PUT /api/connectors/{provider}` — create or replace one connector.
- `DELETE /api/connectors/{provider}` — disconnect one provider.
- `GET /api/series?days=30` — authenticated canonical provider envelopes. Cost windows use UTC calendar dates, not the last N reported rows: `period_start` is inclusive and `period_end` is today (any reported partial day is retained). Sparse and missing observations do not extend the window; missing dates are not confirmed zero spend. The AWS envelope also includes regional EC2 inventory refreshed independently of billing snapshots (at most 60 seconds cached). `?refresh=1` checks EC2 immediately without rerunning Cost Explorer. `inventory_generated_at` records the AWS check time; failed checks return `instances_available: false` and an explicit error instead of old running states. Each instance includes `compute_estimate` with monthly USD (730 running hours), hourly USD, and an explicit reason when unavailable. `instance_pricing` reports the AWS public Linux shared-tenancy On-Demand price publication date and source. These compute-only comparisons exclude storage, snapshots, public IPv4, data transfer, CPU credits, taxes, and discounts; stopped instances show zero compute while other charges may continue.
- `GET /api/series/{provider}?days=30` — query one provider independently.
- `GET /api/projection?days=30` — read the scheduled AI forecast with fresh owner-scoped BrightWrapper recorded costs overlaid as a floor. This read never generates an AI request; the recorded-cost fetch is bounded to ten seconds and cached for 60 seconds.
- `POST /api/metrics/query` — run a bounded AWS CloudWatch metric query.
- `GET /api/fleet-health` — every running EC2 instance in the connected account's region with 24 hours of 5-minute CloudWatch history in one batched read: CPU average/peak, CPU credit and surplus balances, charged surplus credits (`cpu_charged`, AWS `CPUSurplusCreditsCharged`, `Sum`), network out, EBS read ops and I/O burst balance, plus memory and root-disk use where the CloudWatch agent publishes them (`null` means no agent). The table shows the CPU mean over the window, current borrowed credits, and the sum of charged credits over the window. Borrowed surplus is repayable and does not prove a charge; missing charge datapoints are unavailable, not zero. Zero earned credits alone does not prove throttling. `GET /api/server-health?resource_id=<instance>` opens the detailed chart for any instance in that account.
- `GET /api/me` — authenticated viewer identity.
- `GET /api/health` and `GET /api/version` — public operational checks.

System health records daily account-wide AWS billing API charges and live Jessald app-box disk usage. AWS billing is imported at most once every six hours; Refresh checks live EC2 inventory without buying another billing query. AWS `daily_service_costs` and `service_breakdown` cover the same UTC dates as `daily_costs` and the total, including revised historical charges. `billing_observed_at` identifies the last successful billing read.

Connector configuration is encrypted at rest, scoped to the authenticated user,
and never returned by the API. An agent may retrieve credentials from the user's
preferred secrets manager, then send them directly over HTTPS; Observatory does
not need access to that secrets manager. `PUT` is idempotent and replaces the
previous credential set.

Authenticated calls accept either the user's AuthReturn JWT or that user's
app-specific AuthReturn API key in `Authorization: Bearer …`. Agents should use
an API key for unattended access; keys remain user-scoped and revocable.

## Connect a provider

1. Ask the user which provider to connect and which secrets-manager entry holds
   its dedicated billing credential. Read the secret directly from that manager;
   do not ask the user to paste it into chat.
2. Authenticate as that same Observatory user. For an interactive, bounded setup,
   use the user's Observatory AuthReturn JWT. For unattended refreshes, use that
   user's app-specific, revocable Observatory API key. Never use another user's
   bearer credential.
3. Call `GET /api/connectors` and select one of the provider IDs below. Send the
   exact required fields in a `PUT /api/connectors/{provider}` request over HTTPS.
4. Confirm that the response reports `connected: true`, then call
   `GET /api/series/{provider}?days=30`. Report provider permission errors as-is;
   do not broaden the credential automatically.

For a bounded setup, obtain the JWT from the canonical app-scoped login endpoint.
Read the email and password from the user's secrets manager and keep the returned
token out of logs:

```http
POST https://authreturn.com/api/apps/observatory/login
Content-Type: application/json

{"email":"<Observatory account email>","password":"<Observatory password>"}
```

Use the response's `token` only as `Authorization: Bearer <token>` on Observatory
requests. Authentication must terminate with a visible error within 20 seconds;
do not retry invalid credentials. For unattended use, store a user-scoped
Observatory API key in the user's secrets manager and revoke it when the agent
no longer needs access.

Example request:

```http
PUT /api/connectors/openai
Authorization: Bearer <user JWT or app-specific API key>
Content-Type: application/json

{"config":{"admin_key":"<OpenAI organization admin key>"}}
```

The authenticated connector endpoints are on
`https://observatory.aisloppy.com`. Apply an explicit timeout to every request;
30 seconds is suitable for connector configuration and 45 seconds for the first
provider query. The API never returns the submitted credential.

Supported provider configuration:

| Provider | Required fields | Optional fields |
|---|---|---|
| `aws` | `access_key_id`, `secret_access_key` | `session_token`, `region` |
| `azure` | `tenant_id`, `client_id`, `client_secret`, `subscription_id` | — |
| `gcp` | `project_id`, `dataset_id`, `table_id`, `service_account` object | — |
| `openai` | `admin_key` | — |
| `anthropic` | `admin_key` | — |
| `scrapingbee` | `api_key` | — |

## Connect AWS costs and CloudWatch

The existing `aws` connector powers both Cost Explorer and CloudWatch. Use a
dedicated IAM principal with this minimum read-only policy; scope resources
further when the account's metric naming allows it:

```json
{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Action":["ce:GetCostAndUsage","cloudwatch:GetMetricStatistics","cloudwatch:GetMetricData","cloudwatch:ListMetrics","ec2:DescribeInstances"],"Resource":"*"}]}
```

The metrics endpoint accepts JSON containing `namespace`, `metric_name`, an
optional `dimensions` object, `statistic`, `period`, and `hours`. Queries are
limited to 30 days, 1,440 points, ten dimensions, known CloudWatch statistics,
and a 20-second overall deadline. Successful identical queries are cached for
60 seconds. Example:

```http
POST /api/metrics/query
Authorization: Bearer <user JWT or app-specific API key>
Content-Type: application/json

{"namespace":"AWS/EC2","metric_name":"CPUUtilization","dimensions":{"InstanceId":"i-0123456789abcdef0"},"statistic":"Average","period":300,"hours":24}
```

CloudWatch errors are visible but connector credentials are redacted. The API
returns timestamp/value points and never returns the AWS credential.

Provider-specific payload shapes:

```json
{"config":{"access_key_id":"<AWS access key ID>","secret_access_key":"<AWS secret access key>","region":"us-east-1"}}
{"config":{"tenant_id":"<Entra tenant GUID>","client_id":"<app GUID>","client_secret":"<client secret>","subscription_id":"<subscription GUID>"}}
{"config":{"project_id":"<GCP project>","dataset_id":"<billing export dataset>","table_id":"<billing export table>","service_account":{"type":"service_account","project_id":"<GCP project>","private_key_id":"<key ID>","private_key":"<private key>","client_email":"<service account email>","client_id":"<client ID>","token_uri":"https://oauth2.googleapis.com/token"}}}
{"config":{"admin_key":"<OpenAI organization admin key>"}}
{"config":{"admin_key":"<Anthropic organization admin key>"}}
{"config":{"api_key":"<ScrapingBee API key>"}}
```

These examples are ordered as AWS, Azure, GCP, OpenAI, Anthropic, and ScrapingBee.
The Azure credential is a service principal with only the Cost Management Reader role on the subscription.
Send only the payload for the selected provider. Optional AWS `session_token`
credentials expire; replace the connector before expiry or use a dedicated
long-lived read-only IAM principal.

Use read-only billing or usage credentials with the narrowest provider permissions
available. Do not put credentials in query strings, logs, chat messages, or source
control. After configuration, query `/api/series/{provider}` to verify access; a
failed provider returns a visible error without affecting other connectors.

## Security model

- Public signup is enabled. Every connector, provider cache entry, and usage-history
  path is scoped by the verified AuthReturn user ID.
- Connector secrets terminate at Observatory: the service must decrypt them in
  memory to call providers. This is encrypted-at-rest isolation, not a zero-knowledge
  vault. Host administrators retain infrastructure-level access.
- Credentials are never returned after configuration. Sensitive values are redacted
  from upstream errors, protected API responses are `no-store`, and rendered provider
  content uses text nodes rather than HTML injection.
- Browser sessions require an AuthReturn Cognito ID token with verified issuer,
  audience, signature, expiry, and token type. Automation keys are verified against
  the `observatory` AuthReturn app and successful verification is cached for at most
  60 seconds.
- Use a dedicated read-only billing identity per provider. Do not supply general
  cloud-administrator, resource-write, or payment-method credentials.

Minimum credential guidance:

| Provider | Recommended boundary | Residual power |
|---|---|---|
| AWS | A dedicated IAM principal allowing `ce:GetCostAndUsage`, `cloudwatch:GetMetricStatistics`, and `cloudwatch:ListMetrics` | Do not reuse an infrastructure automation key. |
| GCP | BigQuery Job User plus Data Viewer limited to the billing-export dataset | The service-account key remains a reusable bearer credential. |
| OpenAI | A dedicated organization admin key used only for usage and costs | OpenAI's usage/cost endpoints require organization-level admin authority. |
| Anthropic | A dedicated organization admin key | The usage endpoint requires admin authority. |
| ScrapingBee | A dedicated account key | ScrapingBee keys can make paid scraping requests; no read-only usage key is assumed. |

Observatory currently accepts encrypted inline credentials; it does not yet
resolve references from an external user-controlled vault or use provider OAuth.

## Connector adoption analytics

Observatory tracks its own usage under the preserved app slug `observatory`. The backend emits
these privacy-safe events:

- `connection_attempted`: an authenticated connector `PUT` reached the backend.
  This is the primary adoption KPI.
- `connection_succeeded`: connector configuration passed validation and was
  encrypted and stored. It does not claim that the upstream provider accepted it.
- `connection_failed`: the request ended in one of `invalid_request`,
  `unknown_provider`, `invalid_config`, or `storage_error`.
- `guide_opened`: a signed-in user followed a disconnected card's setup link.

Event paths contain only the event, a catalog provider ID, and the controlled
failure category. Attempt sessions use an app-secret HMAC of the AuthReturn user
ID so usage analytics can count unique attempting tenants without receiving the raw
ID or email. Connector configs, API keys, secrets, raw errors, request bodies,
and email addresses are never analytics fields.

View the `observatory` panel at
`https://gridglance.com/`. Attempt totals and provider
breakdown are the page counts under `/events/connection_attempted/{provider}`;
attempt-only session analytics provide the unique-tenant count. No separate
Observatory analytics dashboard is maintained.

The cost-series response is schema version 2. AWS, GCP, OpenAI, Anthropic, and
ScrapingBee return independent availability envelopes; missing or failed
providers are never silently replaced with zeroes. Windows are bounded to
2–366 days. The dashboard includes a stacked daily cost graph, provider totals,
token usage, service/model breakdowns, and ScrapingBee credit headroom.

The final chart date may use a BrightWrapper projection because provider billing
buckets publish at different times. Projection payloads are labeled
`kind: ai_projection`, retain each provider's `observed_usd`, and never change
provider totals or cached source observations. Recent non-read-only CloudTrail
events are weak context when the connected AWS identity permits lookup; an
unavailable CloudTrail query is explicit and does not block history-based estimates.

Base URL: `https://observatory.aisloppy.com`

## Guide routes

`/welcome` is the minimal public landing page, without app navigation. Signed-out
visitors to `/` go there; signing in returns to the existing costs dashboard or
the original same-origin deep link. Signed-in visitors can open the app directly.

`GET /agent-guide.md` is this machine-readable contract; `/agent-guide` is its
rendered page. All browser pages are public shells; their account data remains
authenticated. The supported navigation is Costs & infrastructure, System health,
Storage, Computer, Usage, Profiles, and Agent guide.

## Reviewed peak investigations

`GET /api/server-health/investigation` returns `{available, data}` for the authenticated user’s currently monitored AWS resource. A report is a dated snapshot with separate CPU (%) and memory (GiB) series, overview and detail observations, event annotations, evidence excerpts, and explicit attribution limits. Missing reports return `available: false`. These historical snapshots are available through the API only; the dashboard shows current fleet health. Reviewed reports are stored in Observatory’s tenant-scoped `peak_investigations` table; they are not inferred automatically from timing correlations.


## Structured infrastructure embeds

`GET /api/embeds/infrastructure` returns a public HTML renderer, never tenant data.
Embed this exact URL in an iframe with `sandbox="allow-scripts"` and
`referrerpolicy="no-referrer"`. Only HTTPS FairyStack origins may frame it.
The response's CSP permits only hashed renderer scripts/styles and prohibits
network requests, forms, other frames, and same-origin privileges. It accepts no
HTML, script, credential, or caller-supplied URL.

On iframe load, transfer one MessageChannel port using
`postMessage({type: 'observatory:connect'}, '*', [port2])`. The wildcard is needed
for the opaque sandbox origin; transfer only the port, not evidence or tokens.
The renderer checks the sending window and FairyStack origin, binds once, and
returns `{type: 'ready'}` on the private channel. Send the report on port1 only
after readiness. It replies `{type: 'rendered'}` or
`{type: 'failed', error: '...'}`. The host must terminate a missing readiness or
render acknowledgement within 12 seconds and close the channel on removal.

Reports have `schema_version: 1`, `source_id`, `label`, `status` (`observing` or
`action_required`), `checked_at` (Unix seconds), `samples` (at most 1440), and
`annotations` (at most 82). Each ordered sample has `timestamp`, `load` (5-minute
load average), `cpus` (positive integer), `memory` and `disk` (used percentages).
Unavailable metrics are explicitly null. An annotation has `timestamp` and a
plain-text `label` (at most 300 characters). Invalid reports fail visibly.
Load is graphed per CPU, separately from capacity percentages. Missing samples
and collection gaps over three minutes break the line. Zoom, selected evidence,
focus, and unchanged nodes survive report updates.

FairyStack owns pair-health collection, bounded history, and incident delivery in
this first integration; Observatory owns the renderer. The embed does not store
reports or connect an Observatory user's account. Monitoring migration requires
a separate authenticated connector; this endpoint does not grant that access.


### Warning history and coverage

The default embed window includes both recorded warnings and sampled measurements.
`alerts` is an optional array (at most 200) of archived mailbox observations with
`id`, Unix-second `timestamp`, `kind: recorded_alert`, plain-text `label` (300
characters), `detail` (4000 characters), and nullable `load` / `cpus` values.
These measurements are isolated points, never interpolated into sampled history.
`unattributed_alerts` has the same shape but is explicitly labeled source unknown
and never plotted as a source's load measurement. An annotation may also preserve
its recorded numeric `load` and positive integer `cpus` after samples expire.
The time axis includes events outside the sampled interval; coverage is stated
above the chart. Warnings are selectable above the plot, with original text in a
shared disclosure. Explicit recent windows may exclude earlier events.

The renderer sends `{type: resize, height: <pixels>}` over the private port when
its content changes. Hosts may adjust the frame height within a bounded range
(350–850 pixels); this grants no navigation or parent DOM access.


## Historical server-health links

`GET /api/server-health?start=<ISO UTC>&end=<ISO UTC>&resource_id=<EC2 ID>`
uses the authenticated user's AWS connector and configured monitored server.
Both bounds are required together. A mismatched resource returns 400; missing
configuration returns 409. Queries have a 25-second overall deadline (504 on
expiry), and per-series missing data or upstream failures remain explicit.
The default remains the latest 24 hours. Historical bounds also work on
`POST /api/metrics/query` as `start` and `end` instead of `hours`; they are
included in cache identity and limited to 30 days and 1,440 points.

An existing server-health chart can open directly on an incident with
`/?health_start=<ISO UTC>&health_end=<ISO UTC>&resource=<EC2 ID>&warnings=<JSON>`.
`warnings` is a URL-encoded array of up to 20 `[Unix seconds, load, CPU count]`
records. These are caller-provided event annotations, not CloudWatch metric
values. The historical view shows the authenticated CloudWatch CPU, RAM, and
disk percentages, with warning times marked separately at the top of the chart.
It verifies server identity and omits unrelated cost panels in this view.


### Authenticated history in infrastructure embeds

A trusted FairyStack parent may set `history_enabled: true` on an infrastructure
report. The renderer requests `{type: 'history-request', key: <warning key>}` on
the existing private port. The parent authenticates its own inbox-history API,
which derives the episode window and EC2 identity from the owned incident.
Reply `{type:'history',key,status:'completed',data:<server-health data>}` or
`{type:'history',key,status:'failed',error:<plain text>}`. Historical data must
include explicit ISO `start`/`end`, a period, and CPU/memory/disk series. The
renderer validates percentages, ordered timestamps and incident coverage.

No credential or new network permission is added to the embed. The parent owns
a 35-second lookup deadline; the renderer has a 37-second missing-reply deadline.
Failures keep the original warning Details accessible. Selecting an older episode
requests its own window; unchanged evidence preserves plot nodes and selection.
The same percentage plot renderer is shared with Observatory's authenticated
historical view. Without `history_enabled`, recorded-load rendering remains
available for older parent versions during coordinated rollout.

OpenAI exposes `usage_today` (UTC date, availability, input/output tokens and
model requests), `observed_at`, and `latest_billing_date`. Current-day usage is
shown independently of billed dollars: absent billing is pending, never proof of
zero spend. Usage counts cover the organization's completions usage endpoint;
they are not credit charges. The legacy static token-price estimates must not be
used to reconcile credit balances. Source contract:
https://platform.openai.com/docs/api-reference/usage

## OpenAI usage freshness

OpenAI's completions usage API supports minute buckets; bucket width alone does
not guarantee ingestion freshness. Its Costs API currently uses daily buckets.
For an expensive running job, BrightWrapper's dashboard shows recorded
per-request costs sooner (Observatory does not duplicate it); final charges for a request may remain unknown until it finishes.
Sources: https://developers.openai.com/api/reference/resources/admin/subresources/organization/subresources/usage/methods/completions
and https://brightwrapper.com/agent-guide.md

## Prepaid provider credit balances

OpenAI and Anthropic cards show prepaid API credit balances separately from
spend, subscription allowances, and BrightWrapper credits. Their connected
public admin APIs do not currently provide a supported remaining-balance read.
Without a confirmed observation, the cards explain that the connected API does
not expose a prepaid credit balance and link directly to the provider's billing
page. “Connected” describes the cost feed, not live credit-balance access.
Absence is never rendered as zero. Refreshing costs cannot retrieve a balance.

An authenticated agent that has actually read the matching provider account's
billing page can record that observation with
`PUT /api/credit-balances/openai` or `/api/credit-balances/anthropic`:
`{"balance_usd":123.45,"observed_at":"2026-09-09T05:00:00Z","account_label":"Organization name"}`.
This is an example payload, not an account balance. Do not infer this value from
usage costs, assume grants or top-ups, or supply an unverified number.

`GET` on the same route returns the observation or explicit unavailable state.
`DELETE` removes the observation. Writes require an existing provider connector,
finite USD amount (negative balances are supported), a timestamp with timezone,
and an account label. Older observations cannot replace newer ones. Records
are user-scoped and bound to the current connector revision; replacing or
removing that connector invalidates the old account's balance. The card always
labels a stored value “last confirmed”, and marks it stale after 15 minutes.
Provider-series responses include `credit_balance` independently of cost caches.

Sources: https://help.openai.com/en/articles/8264644 and
https://support.claude.com/en/articles/8977456-how-do-i-pay-for-my-claude-api-usage

## Disk usage treemap

`/storage` is the authenticated disk monitoring page. Each owner may enroll up to
128 EC2 servers or personal devices without changing the existing CloudWatch health preset or billing
connector. The server links select a fixed instance; unknown or foreign IDs fail
rather than displaying a different server. Existing single-server scans migrate
atomically and remain the default view. Selecting a directory drills down; file
names and directory sizes are private account evidence. There are no deletion controls.

### Enrollment and collector credentials

Owner authentication is the normal Observatory JWT or app-specific AuthReturn key.
Never copy that account credential to a collector.

- `GET /api/storage-resources`: `{resources:[{resource_id,label,provider,updated_at,collector_enabled?}]}`.
  The existing health preset is included even before explicit storage enrollment.
- `POST /api/storage-resources`: `{resource_id:"i-…",label:"Finance app box"}`;
  idempotently enrolls or renames one resource, returning `{resource}`. Personal devices use `device-<24 lowercase hexadecimal FairyStack device ID>` and return `provider:"device"`; existing EC2 resources return `provider:"aws"`. EC2 IDs must
  have 8 or 17 lowercase hexadecimal digits. Labels have 1–80 characters.
- `PUT /api/storage-resources/<resource_id>/collector`: returns `{resource_id,
  revoked:false,ingest_token}`. This **rotates** the server's upload-only key;
  store the response directly in SOPS. The plaintext is shown only once. Lost
  responses require an explicit rotation; this endpoint must not be auto-retried.
- `DELETE /api/storage-resources/<resource_id>/collector`: idempotently revokes
  that server's key immediately. Retained scans remain readable. Missing or
  foreign resources return 404.
- `PUT /api/server-storage/ingest`: accepts only a dedicated `obs_storage_…`
  bearer key. It can upload only that key's owner and storage resource; it may GET `/api/storage-client` for its client updates, but cannot read
  snapshots, enroll/rename servers, rotate credentials, access billing, or execute
  commands. Wrong resource returns 403, invalid/revoked key 401, invalid schema
  422. Keys are stored only as SHA256 hashes. Responses and owner APIs are no-store.
- `GET /api/server-storage?resource_id=<EC2 ID>`: owner-scoped `{available,
  resource,resources,snapshot?,stale?}`. Omit the parameter for the existing health
  preset, or first explicitly enrolled server. Missing scans return
  `available:false`; unconfigured/foreign IDs return 409. Legacy owner-authenticated
  `PUT /api/server-storage` remains supported for existing installations.

Each upload has schema 1, `resource_id`, 32-lowercase-hex `run_id`, epoch-second
`started_at`, `deadline_at` (at most 960 seconds later), and status
`running|completed|partial|timed_out|failed`. Terminal results include
`completed_at`; directory results include `total_bytes`, `used_bytes`,
`available_bytes`, `reserved_bytes`, `scanned_bytes`, `entries`, up to 50,000
`{path,bytes,kind,groups?,leaf?}` nodes including `/`, `granularity_bytes` (1 GiB),
at most 20,000 small-entry group byte counts, and up to 20 `{path,error}` entries.
`kind` is `directory` or `file`. Payloads are at most 16 MiB. Uploads return
`{stored:true}`; obsolete/conflicting runs return `{stored:false}`. Exact terminal
retries are idempotent. Older scans never replace newer scans. A running scan
past its deadline is persisted as `timed_out`; a late worker cannot resurrect it.
Previous directory evidence stays available and visibly labeled while a scan runs
or fails. No received scan and a scan older than two hours remain explicit.

### Install on an explicitly authorized server

The existing `scripts/collect-storage.py` scans allocated blocks on the root
filesystem, without following symlinks or crossing devices. Hard-linked files
count once. It sends directory names, large-file names and sizes, **never file
contents**. Names may still be sensitive: enrollment requires the owner's consent.
Directories and individual files of at least 1 GiB are retained; immediate smaller
directories are named leaves, and smaller files/metadata are grouped. Missing
bytes can include filesystem metadata, open deleted files and scan gaps.

Persist `{resource_id,ingest_token}` in host-specific SOPS
`hosts/<host>/observatory-storage.json`, then run the canonical secrets sync.
Run `sudo python3 scripts/install-storage-collector.py --resource <EC2 ID>` from a
verified copy of the committed scripts. The installer verifies local IMDSv2
identity and the matching scoped secret, provisions a root-owned 0600 consumer,
and enables the hourly scan and minute reading timers. It refuses to replace an active collector. Record
the deployed commit alongside the installation receipt. No inbound listener,
SSH credential, AWS account key or general Observatory key is installed.

The collector runs with a read-only filesystem, one read-only secret mount,
only `CAP_DAC_READ_SEARCH`, low I/O priority, 25% CPU quota and a 512-MiB memory
limit. Root-level metadata access is needed for complete sizes; the script is
root-owned. HTTPS uploads refuse redirects. Idempotent uploads retry connection failures and HTTP 502/503/504 up to three times within a 20-second total deadline (at most 15 seconds per call); authentication, validation, and other failures are terminal. The scan
has a 900-second deadline, the process a 960-second alarm, and systemd a 990-second
limit. Scan errors become partial/failed results; lost workers time out in the
server view. The timer deduplicates scans through one systemd unit. Browser reads
terminate in 15 seconds, pause observation while hidden and stop at terminal state.

### Computer view

`/computer` frames the Computer diagram (computer.jessald.fairystack.com `?embed=1`)
for any instance in the connected AWS account. `GET /api/computer` (authenticated,
read-only, 15 s cache, 25 s deadline) returns `{servers:[{instance_id,name,state,
instance_type,collector}],machines:[…]}`. Each machine carries CloudWatch
`cpu_percent` (latest 5-minute average) and `cpu_count` from the instance type for
every server; memory, swap activity, load, filesystems, disk I/O, network, ports,
top processes and systemd services come only from that server's own collector and
are absent otherwise, which the diagram shows as not reported. Readings older than
five minutes are labelled stale. The enrolled collector pushes one ~1 s sample every
minute with `collect-storage.py --reading` (`observatory-reading.timer`) to
`PUT /api/server-vitals` `{schema_version:1,resource_id,reading}`, using the same
scoped `obs_storage_…` key (or the account's agent key for the primary resource);
only the latest reading per server is kept. Both installers install that timer.

`GET /api/storage-health` returns owner-scoped UTC daily counts for the last 30
days: attempts, terminal, failed, timed_out, partial, completed, and max_seconds.
Compact run receipts are retained for 90 days; recovery never erases an earlier
failure. Existing current/previous snapshots seed the history once, without
inventing older attempts. Reading health also terminates expired workers through
the snapshot timeout owner. The daily `collect_system_health.py` job records scan
failure rate and longest scan duration in System health; `--storage-only` refreshes
just these series. Collector upload keys cannot read this endpoint.

### Mac devices through FairyStack

Use the owner's existing FairyStack `/api/companions` connection to install
`scripts/install-storage-mac.py`; no new desktop app, inbound port, administrator
access, or account credential is needed. Enroll `device-<paired device ID>` using
the owner API, then issue its dedicated collector key. Pairing itself never enrolls
a laptop: the owner must authorize storage telemetry (directory names are private).

The installer receives JSON on stdin with `resource_id`, `origin`, `private_key`,
`encrypted_token` and the committed `client_source` (`scripts/storage-client.py`).
Generate an ephemeral RSA-2048 private key on the Mac in a mode-0700 setup directory;
return only its public key. Encrypt the scoped token using RSA OAEP SHA-256 on the
agent host. Only ciphertext may enter FairyStack command receipts. The installer
decrypts on the Mac, saves the key in login Keychain under
`com.fairystack.observatory-storage.<resource_id>`, verifies read-back, and removes
the ephemeral private key after successful installation. Keychain writes use stdin,
not process arguments. No token is persisted in a plist, config, script or log.

The user's LaunchAgent runs on login and hourly while awake. Its mode-0700 working
directory is `Library/Application Support/FairyStack/Observatory/<resource_id>`;
config contains no credential. It scans `/System/Volumes/Data` under the user's
normal rights without elevating privileges. Install Homebrew GNU coreutils with
`brew install coreutils` before installing the Mac collector. It uses `gdu` from
`/opt/homebrew/bin` or `/usr/local/bin` with NUL-terminated records, preserving
newlines, tabs and other filename characters without ambiguity. Native traversal
avoids millions of Python stat calls; hardlinks still count once. Empty directories
can be grouped with metadata. A missing dependency is an explicit failed scan.
macOS privacy denials yield partial
scans; the UI never presents them as full-disk coverage. Sleeping/offline laptops
retain the last observation and show stale after two hours. No file contents are
read or uploaded. This client does not delete files or change cloud sync settings.

For protected app data, the Mac owner must enable the collector's Python executable
in System Settings → Privacy & Security → Full Disk Access. macOS does not provide
a command to grant that permission. Resolve the installed Python executable before
adding it; an interpreter upgrade may change its path. Do not edit the privacy database
or disable platform protections. A scan after the grant verifies actual coverage.

`GET /api/storage-client` accepts only a live storage collector bearer and returns
`{schema_version:1,version,files:{"collect-storage.py":{source,sha256},
"storage-client.py":{source,sha256}}}`. It exposes only these client files. Each
hourly run verifies HTTPS (redirects refused), payload shape, SHA256 and Python
syntax before atomic file replacement. The new scanner runs immediately; updated
launcher code runs on the next hourly invocation. No update failure silently falls
back to old code: it records a failed observation when upload is available. Scans
include `scan_root`, `collector_version` and `collector_revision` (scanner SHA256),
visible in the page's Collector details. No perpetual process must restart to update.

One file lock prevents overlapping scans; the installer refuses replacement of an
active process. The client has a 990-second overall deadline, scans have the existing
900-second deadline, and Keychain operations have 15-second timeouts. Recurring client updates and uploads share the bounded 20-second transient-retry policy; initial installation has a 15-second download timeout. Durable
local `status.json` records the current step and terminal state. A lost collector
uses the existing server-side 960-second timeout. Never-resolving uploads cannot
leave the page indefinitely running. Revoke with the normal collector DELETE API;
uninstall by booting out the matching user LaunchAgent, deleting its plist and
Keychain item. Retained Observatory observations remain readable.

### Resource links

- `https://observatory.aisloppy.com/storage` — current account's default server.
- `https://observatory.aisloppy.com/storage?resource=i-05938f43701e1eef7` — Finance
  app box, visible only to the account that enrolled it.
- `/storage?resource=<EC2 ID>` — IDs returned by `/api/storage-resources`; ordinary
  owner authentication is required. These links select current evidence, not an
  immutable snapshot; the scan timestamp/run ID identifies the observation.

Fleet rollout should reuse this collector and enrollment API with a separate
owner/resource key for every host. This release does not automatically enroll
other customer instances or centralize their filenames under an administrator.
Existing FairyStack capacity alerts remain owned by FairyStack.

## Python release dependencies

Releases install `requirements.lock` with FairyStack’s protected uv wheel cache.
The lock records the complete installed package versions and SHA256 hashes.
Development `.venv` edits do not enter releases. Dependency updates must regenerate
the lock and pass the release import check before activation. Release environments
are read-only; identical wheels share storage across retained releases.

## Unified analytics (v5)

Observatory combines costs/infrastructure, usage, and telemetry. Account data is
isolated by verified Observatory identity. Usage collection and history are native to Observatory. Telemetry profiling history
can still be connected through `PUT /api/connectors/telemetry`, with
`{"config":{"api_key":"ar_…"}}`. Credentials are encrypted per account. Source reads have an
8-second overall deadline and report errors rather than fabricated zero totals.

### Native collection

- `POST /api/analytics/projects`: authenticated owner, body
  `{"name":"my-service","origins":["https://my-service.example"]}`. Returns
  `{id,name,origins,ingest_token}` with a **write-only key shown once**. Duplicate
  names return 409. Keep keys in SOPS; never expose them in browser code.
- `GET /api/analytics/projects`: owned and explicitly shared projects as `{projects:[{id,name,label,origins,access}]}`; read access does not grant ingestion or management authority.
- `POST /api/analytics/otlp/v1/{traces|logs|metrics}`: Bearer project ingestion key,
  OTLP/HTTP JSON or binary protobuf, optional gzip, max 2 MiB decoded / 2048 records.
  Traces retain measured start/end, parent identity, attributes and status; logs
  retain structured bodies; metrics retain type, unit, temporality and data point.
  Batches commit atomically. Exact retries are idempotent; changed immutable span
  identity is 409. Invalid encoding/schema is 422; unavailable storage is 503.
- `POST /api/analytics/browser/{project}`: compatible Faro envelope, max 64 KiB;
  only exact enrolled HTTPS Origin values accepted, 120 batches/IP/minute. This
  public endpoint stores sanitized browser signals and cannot read or write
  server signals. Exception messages and arbitrary event attributes are omitted.
- `GET /api/analytics/observations?project={id}&kind=span&hours=24&limit=200`:
  owner or explicitly authorized viewer; kinds `release`, `preparation`, `span`, `log`, `metric`, `event`, `load`, `failure`; hours 1–720;
  limit 1–1000. Optional `trace={32 lowercase hex}` reads all retained spans for
  the trace independent of the time filter. When `truncated` is true, pass the
  returned `next_cursor` as `cursor` with the same query filters to continue a
  stable time window. Returns records, truncated, next_cursor, window_end, limit,
  retention_days. Retention is 30 days. General traffic is capped at 100,000 records/project; release and preparation traces have a separate 100,000-record budget. Their oldest whole traces are evicted together, so noisy logs and metrics cannot displace release history.
- `GET /api/analytics/sources/telemetry` reads the current account's connected profiling source through its public API. It defaults to the latest 100 profiles; `limit=50|100|250|1000` and exact `project`, `kind`, or `id` filters are sent upstream before reading. `requested_limit` states the selected bound. This compatibility API retains older external evidence for agents; it is no longer called by the Profiles page. The former Grid Glance source route redirects to native `/api/usage/overview`.

### Request profiles

Profiles uses the existing native observation/trace reader at
`/analytics?kind=span`. It reads the selected authorized project's newest 200
spans in the chosen window; truncation stays explicit. Select a chart sample or
row to open the full trace (up to 1,000 retained spans) with measured start offsets
and durations. This includes instrumented HTTP endpoints and dependencies, not
automatic sampling of uninstrumented code. `/profiles` redirects here. It has no
runtime dependency on the older Telemetry collector or its connector credential.

### Resource links

| Browser route | Identity / authorization |
|---|---|
| `/analytics?project={id}&kind=release` | Project ID from projects API; owner login required for records |
| `/analytics?project={id}&kind=span&trace={trace_id}` | Immutable recorded trace; 32-character hexadecimal trace ID; owner-only |
| `/profiles` → `/analytics?kind=span` | Current account-authorized request and dependency traces |
| `/` | Existing costs and infrastructure analytics |

Example trace URL: `https://observatory.aisloppy.com/analytics?project=PROJECT_ID&kind=span&trace=TRACE_ID`.
IDs in this template are placeholders, not existing evidence. Pages are public
shells; all account evidence requires authentication. Missing/expired evidence
renders an empty state. Views refresh on demand, preserve local filters and
selected details, and terminate failed reads within 12 seconds.

Histograms are shown with their recorded sum/count and type in the detail view;
they are not silently treated as individual duration samples. This is a bounded
collector and evidence browser, not a PromQL or LogQL query engine.

Owner-only `POST /api/analytics/projects/{id}/quarantine` accepts
`{"trace_ids":["32-character-hex-id"],"reason":"Proven test contamination"}`.
It excludes those exact traces and their logs from every ordinary read, including
trace links and later replays. Evidence remains retained under normal retention;
the quarantine keeps its reason and timestamp. This is an agent maintenance API,
not a revision-pattern filter: never quarantine a real release by guessed naming.

### Release stage chart

`/analytics?project={id}&kind=release` presents a single stacked column chart,
chronological by completion, one column per attempt (repeated revisions are not
collapsed). It restores the release categories: backend tests, frontend tests,
preparation, signing/delivery, customer activation, marketing activation, queue,
and unattributed time. Colors stay fixed across refresh and queue visibility.
Values are minutes; the queue checkbox subtracts only recorded queue phases.
Columns with missing or inconsistent stage evidence are marked `?`, not inferred.
Hover/focus/tap uses Info Elements’ shared anchored-tooltip component, loaded with
an 8-second deadline. Selecting a column opens the existing measured trace.

Owner-only `GET /api/analytics/releases?project={id}&hours=720&limit=1000` returns the
ordinary release envelope plus each record's `phases:[{name,duration_ms,started_at}]`
from retained OTLP stage spans. Tenant boundaries and quarantine apply. Phase
spans are fetched together, not through one request per chart column. Releases default to 720 hours (30 days). When `truncated` is true, send the returned `next_cursor` with the same project/hours/limit to read the next page; cursors preserve the initial window end and are scoped to the query. The chart reads all pages within a 30-second deadline. Changing Window resets chart zoom; refresh within the same window preserves it. The URL retains the selected hours when opening or closing a trace.

Release speed cards compare the mean backend-test time and total release time of the latest 10 completed, measured releases against the preceding 10 in the loaded window. Total release always includes queues, independently of the chart toggle. Failed attempts and missing stage evidence are excluded from this comparison; fewer than 20 measured completions shows an explicit insufficient-history state.

Release durations default to hourly mean stacked bars, with flush edges and no borders. Average offers individual releases, 1 hour, 6 hours or 1 day. Buckets use fixed UTC boundaries, labeled in local time. Each stage uses the same denominator: releases with complete stage evidence, including failed attempts; missing evidence is excluded and counted in the sidebar. Include gaps defaults on: empty intervals remain blank rather than becoming zero-duration releases. Turning it off compresses only empty intervals; recorded attempts with missing evidence remain visible. Averages and their denominators are unchanged. Interval details include outcome counts and Show individual releases drills into the exact bucket. Latest 10 shows individual releases. Info Elements owns
both drag-range gestures (`chart-range.js`) and hover/focus/pinned details
(`chart-sidebar.js`, `<chart-detail-sidebar>`). The shared components load under
an 8-second deadline and fail visibly; they are not replaced with local copies.
Drag horizontally to zoom; Reset zoom and Latest 10 provide keyboard alternatives.
Refresh defers chart data changes during a drag and preserves the chosen range
and pinned sidebar. The sidebar's Open measured trace opens the existing trace
view. Missing evidence leaves its attempt unfilled and marked ?, rather than inventing zero-duration stages.

The release sidebar passes each category’s canonical color through Info Elements’
`metrics: {label: {value, color}}` contract; swatches match the connected bands and legend.

### Test-run evidence

Analytics storage uses SQLite WAL so readers retain a snapshot without blocking
writer commits. Ingestion applies the same retention limits with covering-index
counts and trims only excess rows; protected release traces are ranked only when
their dedicated cap is exceeded. Write-lock contention has a one-second bound and
returns HTTP 503 `Analytics database is busy; retry the request`; internal logs
retain the SQLite cause. Reads retain their four-second deadline.

`POST /api/analytics/test-reports` accepts a project write-only ingestion key and
an immutable envelope `{id,revision,trace_id,kind,report}`. `id` and `trace_id` are
32 lowercase hex characters, revision is 40; `kind` is publisher or background.
Publisher reports require their release trace. The report schema is version 1,
with gate/full suite, terminal status, updated_at, collected/recorded counts,
elapsed_ms, unaccounted_ms, optional cpu_ms, optional budget `{seconds,exhausted}`
(positive seconds and boolean exhausted), tests (50 slowest with setup/call/
teardown/duration and outcome), files (1,000 aggregates), optional results (10,000
complete per-test outcomes/durations). A completed budgeted run may have fewer
recorded than collected tests: these were not run, not passed or failed. The UI
shows the test-start budget and the unrun count; in-flight tests finish after
the budget, so elapsed time may exceed it. No extra pass circumvents the budget.
Max request 3 MiB. Unknown fields are
excluded. Retries are idempotent; conflicting identity returns 409. Retention is
30 days or 200 reports per project.

Owner-only `GET /api/analytics/test-reports?project={id}&revision={sha}&trace={trace}`
returns `{reports,retention_days}` for the exact publisher attempt and up to two
recent background runs for that revision. Missing reports are unavailable, never
implicitly passed. Quarantined release traces remain excluded.

Resource link: `/tests?project={project_id}&revision={40hex}&trace={32hex}` opens
the per-file and per-test drill-down. Same owner authorization as release traces;
public page shell does not grant access. The Backend tests sidebar metric links
here. Run choice, search, outcome and file filters are local drafts preserved
on refresh. Source reads terminate within 12 seconds.

Release charts default to individual releases, with Include gaps and Include queue wait unchecked. Optional hourly/daily intervals
show means across their attempts, which can differ from a selected release's
trace. Chart categories and measured trace stages share the same display labels;
combined categories retain their specific substep after the category name.

### Signing and delivery drill-down
`/analytics?project=<24hex>&kind=release&trace=<32hex>&stage=delivery` focuses the selected release's retained delivery intervals. Linked from the release sidebar. Owner authentication is required; immutable trace IDs select the attempt. Example: `/analytics?project=01badc780e971908930f965e&kind=release&trace=<release-trace-id>&stage=delivery`. Current publishers measure bundle creation/verification, checksum/signing, and upload/acknowledgement per allocation; older evidence has combined notification intervals. Missing time is explicitly unmeasured. Detail spans do not contribute additional time to top-level release stacks.

### Navigation and report discovery
Every full app page uses `/observatory-header.js`, the single `observatory-header` component. The Observatory brand links to `/`; navigation labels, order and active-page styling come from this shared component. Embedded infrastructure incident charts keep their compact embedded title.
`/analytics?panel=tests` and old Preparation listings now open System health.
Individual `/tests?project=<24hex>&revision=<40hex>&trace=<32hex>` reports remain
available through release evidence. `/telemetry` redirects to
`/analytics?health=observatory`: its measurements concern Observatory's own cost
page, not a separate general telemetry product. Administrators see the latest
100 request/server/render samples there, filtered by the selected health window.
`GET /api/performance` retains its administrator-only authorization; the same
samples appear as `cost_performance` in administrators' system-health responses.

### Retained preparation evidence

Preparation has no standalone product page or navigation entry. Old Preparation
URLs open System health. Its existing ingestion and authenticated evidence API
retain operational records for agents; this UI cleanup does not delete history.

Project-key OTLP logs with body `kind: preparation` require operation, pair_id,
attempt_id and terminal outcome (completed/failed/timed_out/cancelled).
Optional duration_ms may be null for absent measurements. Companion spans carry
actual start/end times and parent identity. No private error text or credentials
are exported from preparation receipts.

Native reads use kind/time and kind/trace indexes so timing reads do not scan
unrelated runtime telemetry. SQLite query execution has a four-second deadline;
expiry returns 504 with a concrete retry message. The browser retains its
12-second overall read deadline, deduplicates concurrent initial loads, and
preserves edited filters on refresh.

Preparation records retain their immutable trace identity and distinct retries.


## Usage analytics

Usage analytics (tracker, visits, browser IDs, sessions, KPIs, device resources) moved to Grid Glance on 2026-10-07: https://gridglance.com/agent-guide.md. `https://observatory.aisloppy.com/api/usage/*` forwards there for trackers and integrations still using that URL, and the old `/usage`, `/usage-app`, `/usage-session` and `/resources` pages redirect to gridglance.com.

## Browser app loads and observation charts

Browser load charts live inside each app's System health dashboard at
`/analytics?health=<app-key>`. Enrolled browser projects without a named dashboard
receive one at `/analytics?health=project:<24hex>`. The directory combines the
existing dashboards with the authenticated viewer's enrolled browser projects;
projects matching a named app share that dashboard. No standalone App loads,
Preparation, or Backend tests tab remains. Individual backend test reports are
still available from release evidence; old `panel=tests` URLs open System health.

The chart uses the app's own collectors and the selected health window, capped
at the 30-day navigation retention period and newest 200 samples per collector.
Truncation and missing collectors are explicit. Select a bar for the immutable
navigation's stage/resource timings; `trace=<32hex>` preserves that selection.
Old `kind=load&project=<24hex>` URLs resolve to the corresponding app dashboard.

`GET /api/analytics/observations` still accepts `kind=load`, selecting only
`fairystack_load` browser events, and the same trace API supplies measured spans.
The browser uses the existing bounded analytics reader and ignores stale responses
when the app, window, or selected navigation changes.

The existing origin-restricted browser collector accepts `fairystack_load` events with the existing allowed status, kind (failed/current step), duration_ms and trace_id fields. Its related OTLP browser spans preserve only bounded numeric `load.*` fields: transfer/encoded bytes, resource count/omitted count, connection downlink, explicitly emulated Mbps (zero means no declared emulation), milestone/resource flags, response wait/download milliseconds; plus sanitized status, step and version. No URL queries, API identifiers, message content or credentials are recorded. Static asset names are retained; API requests and cross-origin resources use generic labels.

FairyStack's lightweight recorder starts near the beginning of HTML parsing and uses navigation's time origin. Terminal completion means the restored initial surface is rendered, window load fired and images belonging to that first conversation snapshot settled; it excludes future polls, later images and optional telemetry SDKs. A failed initial image is a failed load, not successful completion. Startup has a 180-second deadline from navigation, page exit records cancellation, and collection has a separate 15-second deadline. A reload is a new trace. Recorder delivery failure is exposed locally in `FairyStackLoadProfile.state()` and console; a lost/unreachable collector cannot promise a retained record. Browser/process termination before delivery can leave a gap. Human login time can be included; the authentication milestones identify it.

Waterfalls contain elapsed-from-navigation milestones and overlapping resource intervals: do not sum them. At most the 60 longest completed resource requests are retained; root fields show omissions. In-flight requests have no final resource duration. Zero transferred bytes may mean a cache hit or cross-origin timing restrictions. `load.downlink_mbps` is a browser estimate, not a measured speed test. Synthetic evidence must declare `load.emulated_mbps`; it is not a measurement of the operator's connection.

## Default app profiling

`PUT /api/analytics/sites/<exact-hostname>` enrolls one browser-only project for
an authenticated Observatory owner. Repeating it returns the same project;
other owners receive separate projects. Response: `{status:"completed",
project_id,origin,viewers,collector_url,script_url,view_url}`. No server ingestion key
is issued. Invalid hostnames return 422, missing authentication 401, busy storage
503. Use a ten-second request timeout.

The optional body `{"viewers":["operator@example.com"]}` (at most 10 emails)
replaces the project's read-only viewers; omitting it keeps them. A viewer signed
in with that verified email sees the project in `GET /api/analytics/projects`
(`access:"viewer"`, with a readable `label`) and reads its observations, releases
and test reports. Quarantine and enrollment remain owner-only; every other
account receives 404 `Telemetry project not found`. FairyStack grants the
installation owner so projects its agent enrolls are readable under the owner's
own login.

The small `/app-profile.js` SDK is copied into each new FairyStack app at creation
and served locally so an Observatory outage cannot block app startup. Include
`<script src="/observatory-profile.js" data-project="PROJECT_ID"
data-completion="manual"></script>` early in the head. It automatically appends
one collapsed, selectable timing tree. `ObservatoryProfile.begin(name,parentId)`
returns a span ID; `end(id,status)` closes it. `start(name)` begins a fresh measured navigation in the same collapsed tracker; `finish(status)` closes it. The tracker defaults on and offers an origin-scoped browser opt-out inside its disclosure. Opt-out survives reload and can be reversed with Enable and reload. Record/export pages omit profiling. `finish(status)` terminates the
initial load after data has rendered; statuses are completed, failed, cancelled
or timed_out. Without manual mode, window load completes the profile. Startup
has a 60-second deadline; collector delivery has a separate five-second deadline
and visible failure state. Reload creates a new trace; page exit cancels pending
startup. Resource intervals are captured automatically. Server-Timing durations appear locally as duration-only bars; their start offsets are unknown and are not exported as measured trace intervals.

Backend spans can be returned in the app's ordinary authenticated API response
and attached with `addSpans(rows,parentId,requestStart)`: rows contain id, parent,
name, start, duration (milliseconds from request start), status. Parent rows must
precede children. Only measured, non-sensitive names and intervals belong here.
No response bodies or URL queries are collected. This is elapsed-time tracing,
not a sampling CPU profiler; uninstrumented server internals remain unknown.
Browser evidence cannot independently authenticate server spans. Bars overlap
and must not be summed. At most 60 supplied spans and 60 resources are included.


Mac release handoffs are measured gaps between named pipeline stages, displayed
separately as “Release handoffs (internal steps unmeasured)”. Their boundaries do
not measure the internal distribution of packaging, dispatch, or waiting. Failed
update-check attempts remain Update verification spans. Release axis labels use
local 12-hour time with AM/PM.


## System health

App-content builder corrections count weekly corrective commits on FairyStack’s
shared release line for compact app surfaces and information overwhelm. This
measures recurring repair work from commit subjects, not app page density or
proof that the default works. The operator judges the trend.

FairyStack regression test lines and product/operations lines are recorded daily
by `scripts/collect_system_health.py` from the shipped Feature ledger counts on
Jessald's `/api/client-diagnostics/health` (`source_size`). These neutral metrics
have no target or ratio: they expose maintenance growth for operator judgment.
Missing source accounting fails collection explicitly rather than recording zero.

Multi's real-user end-to-end check runs daily at 04:30 UTC (up to five minutes
of scheduling jitter). It uses a fresh ordinary account and the actual signup
form and composer, asks the live agent to create one uniquely named static
counter app, verifies its public HTTPS page and button, then archives its own
conversation and deletes its exact test app and account. Cleanup has separate
administrator credentials; those never enter the tested browser. The build has
a 20-minute deadline and cleanup four minutes; systemd provides a hard watchdog.
Interrupted runs are recovered before another signup. The timer schedules one
attempt daily. For an explicit rerun use
`sudo systemctl start --no-block observatory-multi-e2e.service` and follow
`journalctl -fu observatory-multi-e2e.service`; the existing run continues if the
service is already active. Direct runner invocation is rejected. Every service
start performs a real attempt, even after an earlier attempt that day, so a
skipped check cannot mark a failed run healthy. Success including cleanup clears
the active service warning automatically; failure or timeout keeps it visible.
Historical receipts and failure-rate charts retain all attempts. It never
allocates servers.

`GET /api/system-health/multi-e2e` exposes the latest 60 run receipts to the
operator, including the concrete failed step and cleanup state, without passwords
or member data. Existing FairyStack System health charts include failure rate
and age since the latest attempt; `scripts/collect_system_health.py` updates both.
This checks the signup/build journey, not every generated app or model provider.

`GET /api/system-health` returns the authenticated owner's metric catalog and daily/weekly series.
The `fairystack_stranded_send_commands` series records current pending Sends with no accepting turn, including startup Sends parked for a later turn while their agent runs. It comes from FairyStack runtime health's `input_delivery` snapshot; intentional queues are excluded.
Progress/tool order repairs count unexpected browser footer moves with an unchanged
anchor, received over the last 24 hours from Jessald client diagnostics. Owner and
viewport transitions are excluded; old and offline clients remain unobserved.
Mac and Windows self-update failure rates come from their public hosted GitHub
workflow attempts, including earlier failed attempts after a rerun. The collector
reads the latest 100 runs created within 30 days, groups attempts by UTC start date,
and excludes cancelled, skipped and running attempts. Provider and verification
failures count; these rates do not measure failures on customer devices.
Claude login refresh failures, retry attempts and successful model-access recoveries
are collected from Jessald's content-free `/api/agent-runtime/health`. These use
retained events by UTC date; absent days and deleted sessions are unobserved.
Recovery does not imply the requested task completed.
Codex observation disconnections, recoveries and exhausted recovery failures come
from the same endpoint's `codex_recovery.days`, retained as `fairystack_codex_observation_disconnected`,
`fairystack_codex_observation_recovered` and `fairystack_codex_observation_failed`.
Counts describe transport observation; one turn may disconnect repeatedly or keep
working while FairyStack cannot observe it.
The same endpoint's `sqlite.days` records terminal turns lost to database locks
across all runtimes as `fairystack_sqlite_lock_failed`. Seven UTC dates are
collected; today is partial, deleted sessions are excluded, and recovered turns
retain their failure. This is not a count of every failed SQL statement.
Multi Sheriff's public `/api/sheriff/health` supplies UTC daily counts collected
as `fairystack_sheriff_reviews`, `fairystack_sheriff_failure_rate`,
`fairystack_sheriff_deadline_expired` and `fairystack_sheriff_worker_disconnected`.
These distinguish review deadlines from lost owners, including interrupted web
releases. Recorded failures remain after recovery; unreviewed activity and
detection accuracy are not measured. Today is partial. Recorded timing adds daily
maximum review, provider-work and queue-wait seconds; missing historical or
incomplete provider timestamps remain unknown. `--sheriff-only` refreshes these
metrics without collecting unrelated systems.

`PUT /api/system-health/metrics` accepts `{observations:[{metric,period,value,numerator?,denominator?,note?}]}`;
Puppet Sprites `/api/health/embellishment` supplies `puppet_embellishment_capture_failures` , `puppet_embellishment_comparison_reviews`, `puppet_scene_direction_revision_requests`, and `puppet_scene_sound_revision_requests`. The last counts recorded human direction feedback with explicit revision-request intent by UTC recording date; coverage is partial and multiple requests can concern one defect. Repairs retain feedback. The first two count retained A/B Versions reviews by UTC date; retries preserve failures, identical retries may share one review, and frame comparisons are sampled. They do not cover every film export. The daily collector and `--deployments-only` refresh these series. `puppet_audio_listener_control_failures` and `puppet_video_listener_control_failures` count explicitly failed known-input controls from up to 500 recent saved trials, by UTC evidence update date; these count attempts with partial coverage, not unique causes or general model accuracy.
`puppet_timeline_repair_commits` records weekly non-merge Puppet Sprites fix commits mentioning timeline, scrubbing, playhead or needle, from current Git history since 2026-08-17. The daily collector refreshes it; this measures recurring repair work, not observed drag failures. Unlabelled repairs are excluded and the current week is partial.
Puppet Sprites deployment failures are recorded from terminal app-deployment traces as `fairystack_puppet_deployment_failure_rate`. Failed attempts remain after recovery; successful activation does not establish scene correctness or HTTP continuity.
`fairystack_app_deployment_failure_rate` records failed and timed-out terminal deployments across all measured apps on enrolled boxes, including preparation failures. Successful retries retain earlier failures; cancellations and missing traces are excluded. The existing `--deployments-only` collector refreshes both deployment-failure series.
period is an ISO UTC date (Monday for weekly metrics). Observation notes must contain 1–300 characters. Unknown metrics or invalid values return 422;
missing authentication returns 401. The collector records ArrowSplit live-journal p95 backend latency
and backend failure percentage, plus node/edge save and graph navigation p95 latency and server failure percentage, from its public `/api/health` content-free counters. Samples cover at
most 24 hours / 25,000 reads since the current process started. Restarts reset that window;
no traffic produces no observation. Journal reads exclude authentication; graph reads and writes include it. All exclude browser transport.
`DELETE /api/system-health/metrics` accepts `{retractions:[{metric,period,reason}]}`
for 1–100 known metric observations owned by the caller and returns `{deleted}`.
Use it to retract collector artifacts or observations from an abandoned experiment;
a reason is required, invalid input returns 422, and other owners are unaffected.

The collector also records Jessald main-thread stalls over one second and the p95 duration among reported long tasks (at least 50 ms), from `ui_long_tasks` on `/api/client-diagnostics/health`. These measure blocking browser work, including rendering, rather than network wait or total session-open latency. Chromium supports these observations; unsupported, older and offline clients remain unobserved. No received long tasks produces no sample.

Multi's device specification coverage chart counts distinct non-test people with request evidence in fourteen days but no verified device-profile receipt. Its source is `member_activity.people_without_device_specs` on `/api/client-diagnostics/health`. These people remain included in FairyStack's admin device/browser counts using existing authenticated activity; the chart measures unavailable hardware reports rather than missing users.

`browser_responsiveness` on that same endpoint adds foreground timer/paint delays
(at least 250 ms, including Safari), animated session swipe duration/failures,
and offscreen preview network/render timings. The collector records foreground
delay samples over one second, delay p95, swipe p95 and swipe failures. Hidden
time is excluded; samples may share a stall. Missing reports remain unobserved.

Apps directory opens are recorded as `browser_responsiveness.apps_loads` and collected
as `fairystack_apps_load_p95_ms` and `fairystack_apps_load_failures`. Total latency
includes authentication, network, decoding and first paint; the fifteen-second
browser deadline counts as a failure, navigation cancellation does not. Session
diagnostics retain request/render durations and an API trace ID; server traces
separate inventory from connector-registry lookup. Missing clients are unobserved.

Browser startup spans retain numeric `load.hidden_ms` and a bounded native app/version `load.client` (or `browser`), alongside the loaded revision and phase timings. Raw user agents and account identities remain excluded.

Startup attempts add `fairystack_startup_p95_ms` and `fairystack_startup_failures` from the existing client-diagnostics API. Total startup includes observed hidden time; exact client/hash and request/render/paint stages are linked from FairyStack Session diagnostics. Missing startup reports remain unobserved.

The collector records Jessald browser search/navigation p95 duration and failure
percentage from `session_search` on `/api/client-diagnostics/health`. Measurements
include opening the matching message, exclude typing debounce and cancelled
queries, and contain no query text. No reports means missing coverage.

The collector also records page-update delays, failed refresh retries, and maximum
reported delay from Jessald and Multiplayer's `release_updates` diagnostic aggregate.
These cover distinct browser runs and target releases received in 24 hours, including
intentional microphone/selection deferrals. No reports produce a gap, not a healthy
zero. Finance and offline clients are unobserved by this collector.

### DCF valuation quality trends

`/analytics?health=dcf` focuses the existing System health charts on valuation quality.
Read-only 7/30/90-day and all-collected-history windows retain daily sample counts,
definitions and evidence tables. Calendar gaps break the chart instead of implying
unobserved measurements. The existing collector reads DCF `/api/health`
`audit_health.valuation_trends.observations` and upserts retained daily cohorts.

Build completion is execution, not correctness. Agent-reported usable and bug-free
shares exclude missing structured assessments; coverage is shown separately.
Price-distance mean/median use immutable saved valuations and their build-time market
quotes, not current prices or sell-side targets. Closer to price is not a quality goal.
Rebuilds count as separate saved models; selected reviews overrepresent difficult cases.
Legacy prose is not classified into invented bug counts. Today is partial.

DCF trend queries accept `GET /api/system-health?timezone=<IANA zone>`; the focused
view defaults to the browser timezone with an explicit UTC option. Original UTC
observation keys remain unchanged. Their `samples` retain anonymous original event
timestamps and sufficient statistics so local rates, means and medians are recomputed
from events, including DST boundaries, never relabeled from UTC aggregates.
Missing source timestamps are counted explicitly and omitted from local buckets.

The default history begins September 22, 2026, when retained builds became daily
and higher volume (14, 231, 40, 21, 94 attempts on September 22–26). This is a display
boundary, not a claim that every run came from the mover queue; All history restores
earlier observations. No observation is deleted for low volume or a poor result.
Daily bars show each metric’s denominator. Round pixel-sized markers remain round
at any chart aspect ratio. Hover, focus or tap a point/bar for its value and sample.

The headline and solid line exclude the day still open at the latest collection,
even after the wall clock crosses midnight; incomplete collection is not a closed
cohort. Partial values remain listed separately and may be plotted as hollow markers.
Sample size is shown as evidence, not as an automatic statistical significance verdict.

### Voice Feed outcome trends

`/analytics?health=voice` focuses the existing System health charts on voice
plausibility, apparent cutoffs, full browser-pipeline latency, unexplained missing
heartbeats and explicit failure counts. Quality is a text-only proxy, not word
accuracy. Hidden drafts, old recorders and unreported clients do not enter latency
percentiles; coverage has its own chart. UTC daily points update every five minutes
through `voice-health-collector.timer` and `scripts/collect_system_health.py --voice-only`.
Today is accumulating. Evidence notes retain the pinned rubric, judge and observed
capture builds. Gaps remain blank, including the pre-instrumentation period.
Source `/api/voice/quality/health` on Jessald exposes aggregates only; owner-scoped
passages and review failures remain at FairyStack `/api/voice/quality`.

The same view includes BrightWrapper event-loop stall share and largest timer
delay. The collector reads `/api/health` timer observations; coverage is the
current BrightWrapper process over at most 24 hours, resets on restart, and
cannot attribute a stall to an individual transcription.

The daily system-health collector records `system_health_collection_failures`, with failed steps in its evidence. The former direct-database conversation classifier is retired; previously recorded sentiment and reopened-fix observations remain historical evidence.


## Health dashboards by app

Open `/analytics` (also reached by `/system-health`) for the app directory.
`/analytics?health=voice`, `?health=dcf`, `?health=fairystack`,
`?health=arrowsplit`, `?health=annum`, and `?health=observatory` each show one app.
Shared Finance infrastructure is separate at `?health=finance`.
DCF publishes ten charts, starting with build completion. The dashboard removes duplicate and low-value metric charts, the release experiment panel, browser-load panels, forecast-accuracy plots and session-spending panels. Historical observations and ingestion remain retained. AI costs are at DCF `/costs`.
Release evidence and individual test reports remain available; retired Preparation and Backend tests listing URLs open System health.
`GET /api/system-health` includes `dashboards` and a stable `app` key on each
catalog metric. App ownership follows the user workflow, not the service that
emits the measurement. Each chart includes textual recorded evidence, its
measurement definition and source. History and collection remain unchanged.

System health also records daily FairyStack release failure rates and superseded
request counts from the existing immutable release receipts. Failed/timed-out
attempts count as failures; superseded and cancelled requests do not. Cohorts use
UTC request dates and retain evidence of the failed stage. Collection reuses
`GET /api/analytics/releases` with bounded pagination and a 60-second deadline.


### Anthropic reported costs

Anthropic chart values and `billed_cost` come from the organization's
`/v1/organizations/cost_report`, including input, cache, output and tool charges.
Amounts are converted from decimal cents to USD; all pages are imported within
a 90-second deadline. No static model-price fallback is used. The old
`estimated_cost` and incorrectly parsed token counts are retired. Existing
estimated history is replaced across all four windows by a full annual import.
Old estimates are not shown while that import is pending or has failed.

`cost_source`, `cost_note`, `observed_at` and `latest_billing_date` explain the
provenance. Days are UTC; a missing bucket is unreported, not confirmed zero.
Credit purchases, Claude subscriptions and Priority Tier charges are outside
this report. See [Anthropic's contract](https://platform.claude.com/docs/en/manage-claude/usage-cost-api).
## Session spending

`/analytics?health=dcf#spending` embeds session spending in the owner-authenticated DCF dashboard: a daily stack of metered BrightWrapper API charges and subscription-covered agent usage valued at API list prices. Stack height is not a cash bill; subscription fees and direct provider API calls outside BrightWrapper are excluded. Open Session breakdown and select a session to focus its history, or select a day to inspect contributors. The Subscription-covered agents toggle beside the totals remembers its setting. Spending retains its Finance/BrightWrapper coverage and a maximum 30-day UTC window independently of the DCF valuation-cohort controls. Published `?health=spending` links redirect to the DCF spending section; there is no separate spending dashboard. Untagged calls stay Unattributed API by project; absent prices and stale/failed sources remain visible.

`GET /api/system-health/session-spend` reads or refreshes the owner's source snapshots (ten-minute reuse; the existing scheduler refreshes every 15 minutes). `?refresh=1` explicitly refreshes. Requests have an 11-second source deadline; failed sources retain their last good snapshot, labeled stale/error. `PUT` ingests one `{source,origin,snapshot}` or `{source,status:failed|timed_out,error}` receipt, owner-scoped and idempotent per source. Schema-1 snapshots contain category, observed_at, since, basis and daily rows with session_id, label/project, usd, requests and unpriced_requests. This is a retained 30-day source window, not lifetime costs.

Connect through `PUT /api/connectors/session_spend` with `config` containing `fairystack_sources_json` (JSON string of `{origin,api_key}` for Jessald and/or Finance; dedicated owner keys with only `costs:read`) and `brightwrapper_api_key`. Source keys are encrypted by the existing per-user connector store and never returned. Collection reads only network APIs and exports no prompts. `scripts/collect_system_health.py` also records failed collection steps in the existing collector-health metric.

Private session-spending sources may include `connect_address`, a Tailscale IPv4 address, alongside origin and api_key. This connects over the existing private network while retaining the origin hostname for SNI, certificate verification and Host; TLS verification is mandatory. It is useful where a release-only localhost DNS override otherwise captures the control hostname. No public ingress or release routing changes are required.


### Current-day cost floor

The default chart's UTC current-day point uses the maximum of provider-reported
costs, BrightWrapper recorded costs, and the cached full-day AI estimate.
BrightWrapper's `/api/usage/daily-costs` is read with the owner's existing
`session_spend` connector key, never a shared service account. Only OpenAI and
Anthropic recorded charges supplement their matching cost series; Gemini
charges are not folded into GCP infrastructure billing. BrightWrapper charges
include markup/rounding and overlap provider reports, so they are never added. Fractional cents from the recorded ledger are preserved until dollar amounts are formatted.
They exclude unpriced and unfinished requests; the tooltip identifies recorded
spend separately from the full-day estimate. Historical billing and reported
provider totals remain provider-reported.

Scheduled forecasts also receive today's recorded spend. A stale or missing AI
forecast cannot suppress new recorded spending: `/api/projection` reapplies the
floor on every read and returns `kind: recorded_costs` if no same-day AI forecast
is available. Recorded-cost failures keep today's last good evidence with a
visible stale/error label; no previous-day projection is applied to history.
Refresh reloads the projection as well as the provider snapshots.

### Prospective repair evidence

`GET /api/analytics/observations?project=<id>&kind=failure&hours=1&limit=100`
selects error logs, failed spans, HTTP 5xx and shared app-failure events. Responses
retain bounded `next_cursor` pagination (send `cursor` with the identical filters),
`truncated`, `window_end`, and `feeds` timestamps. Feed timestamps cover recently
received server telemetry or browser error-capture capability/error events; old
load-only browser SDK traffic is not evidence of error coverage. Records received
through current collectors have server-owned `ingestion: browser|server`; callers
cannot choose the ingestion authority. Browser Origin does not authenticate a
caller. Historical records without this marker remain unclassified.

The existing usage tracker loads `/app-errors.js` for pages with an enrolled
`observatory-profile.js` script. Normal page loads adopt current shared error
capture, including apps with older local timing SDKs. It sends fixed codes for
JavaScript exceptions, unexpected promise rejections and same-origin HTTP 5xx;
new timing SDK failed steps use the same reporter. No exception text, request URL,
query, response body or credentials are exported. Capture sends a capability
heartbeat and at most 20 events per page, one at a time with a five-second terminal
deadline. `ObservatoryErrors.state()` exposes completed/failed/timed_out delivery.
No enrolled script means no capture; absence of reports is unobserved, not healthy.


System health also records FairyStack's final-speech Send handoff completions and
failures, plus late page mounts kept away from the active view. The existing
`collect_system_health.py` collector reads `/api/client-diagnostics/health`.
These received-event counts measure delivery ownership; they do not establish
audio accuracy or coverage of offline/old clients.

DCF deployment duration and activation medians are recorded daily by
`scripts/collect_system_health.py` (`--deployments-only` for a focused refresh),
from app-deployments OTLP traces. The existing DCF health dashboard links to the
measured attempt and phase timeline. Activation includes restart and verification,
not measured browser downtime. Missing or truncated evidence never becomes a
zero-duration sample. Initial enrollment and the current UTC day are partial.


## Recorded app cost analysis

Observatory owns cross-app operational analysis: costs, latency, failures and links to execution evidence. BrightWrapper owns its charge ledger; FairyStack owns sessions and their history; apps own domain results. The shared renderer is presentation, not another execution or domain-data owner.

Cache fields on each cost aggregate: `cache_samples`, `cache_input_tokens`, `cache_read_input_tokens`, `cache_creation_input_tokens`, `cache_creation_1h_input_tokens`, `cache_hit_requests`, `cache_savings_cents`. Reads / measured input gives token reuse. Savings compare the measured calls with uncached pricing at charge time, after write premiums and markup; negative means net extra cost. Missing historical/provider measurements remain unknown and are excluded, never treated as misses. The System health collector records today's measured cache reuse and net savings from the same API.

The Costs & infrastructure page begins with the shared recorded-cost explorer.
Deep link: `/?project=dcf-cell-by-cell&days=7&day=2026-09-27`.
Select an app, day or step to see the contributing calls, measured input/output
usage and exact session links. Missing session IDs stay in unattributed groups.
This is separate from provider invoices and subscription-equivalent estimates.

`GET /api/costs` accepts the existing owner-scoped AuthReturn JWT or Observatory
API key. Configure the existing `session_spend` connector; its BrightWrapper key
owns the cost source. Observatory reads BrightWrapper `/api/usage/cost-breakdown`; BrightWrapper owns aggregation. DCF reads that API directly and does not depend on this endpoint. Query: `days=1|3|7|14|30`, optional exact `project`, `day`
(UTC ISO date), `operation`, and optional `related_project` / `related_prefix`
for a separately reported `shared` table. Filters never change owner authority.
Schema 1 returns `totals`, `daily` (whole selected app/window), `steps`, `sessions`,
`projects` (owner's app directory), and `calls` (daily input groups when a step is
selected). Tokens retain sample counts; missing samples remain unknown. Related
charges are not added to the selected app. Session names/links come from retained
owner-authorized FairyStack snapshots; unknown or ambiguous origins stay unlinked.
No prompt/response bodies are returned and reading never makes a paid model call.

Source reads have a ten-second bound and owner/connector-version caches last five
minutes. Failed refreshes return the last successful snapshot with a visible
`error`; an unavailable first read returns 503. API consumers must preserve this
error and observation timestamp. No guessed prices or timing-based attribution.

Reusable browser artifact: `/cost-breakdown.js` and `/cost-breakdown.css`.
`ObservatoryCosts.mount(host, {request(filters, signal), project?, onData?, signal?})`
uses the host's authenticated API adapter, a 15-second overall deadline, aborts
superseded reads and returns `{destroy, reload}`. A fixed `project` hides the app
selector. Hosts may add operation labels and domain evidence in their adapter.
Vendor the published artifacts together. DCF uses the locally vendored renderer but reads BrightWrapper directly, removes the owner's unrelated app list and adds its own review evidence. This component does not require an Observatory data connection.

### Individual system-health charts

Every metric chart title is a permalink. Use `/analytics?health=<app>#<metric_key>` to scroll to and highlight a chart after its data loads; retain query parameters such as `days` and `timezone` to share the same view. For DCF deployment timings, link `/analytics?health=dcf#dcf_deployment_seconds` (total duration) or `/analytics?health=dcf#dcf_deployment_activation_seconds` (activation and verification). These measure deployments, not HTTP downtime.

`fairystack_app_certbot_conflict_rate` reads content-free deployment receipt aggregates from Jessald and Multi. Explicit Certbot busy errors remain after successful retries, including deployments on boxes without telemetry export enrollment. The existing deployment collector refreshes this series on System health.

Unattributed operation rows are selectable. Their recorded input groups show the BrightWrapper dashboard call families (short prompt openings), with their charges and models. These previews explain older work without inventing operation or session IDs; complete-input fingerprint coverage remains separate. New image and speech calls retain their caller-supplied operation IDs.

System health tracks the daily share of BrightWrapper ledger calls with a recorded operation ID. Historical prompt families do not count as recorded attribution; failed and unpriced calls remain in the denominator.

Analytics readers open bounded read-only SQLite connections. Schema and indexes are prepared at serving-process startup; release stage queries use their own measured-phase index so unrelated spans do not enter the chart query. System health records daily median completed release duration and test duration from the same retained receipts.

`puppet_playback_repair_commits` records weekly non-merge Puppet Sprites fixes mentioning playback, rendering or freezing. The daily collector refreshes the repair trend; it does not claim to count browser hangs.

`puppet_sound_repair_commits` records weekly non-merge Puppet Sprites fixes mentioning sound, audio, engine, ambience or SFX. The daily collector refreshes the recurring repair trend; it does not measure missing sounds or judge the mix.


## Shared FPS monitor

FairyStack’s same-origin `/.fairystack-analytics/tracker.js` automatically includes
the current `/app-fps.js` component. No per-app copy, database or extra collector
is needed. Existing apps adopt updates on an ordinary page load. The monitor
defaults on with a compact FPS readout and a collapsed one-minute sparkline.
Its disclosure offers an origin-scoped browser off/on switch, independent of
initial-load profiling. Turning it off cancels its frame loop and persists across
reload; the small “FPS off” control restores it. Hidden pages and record/export
pages do not sample. `<html data-observatory-fps="off">` is an explicit page/app
opt-out across every injected tracker. `data-fps="off"` skips an individual embed;
`data-fps="on"` enables an external tracker integration.

This measures foreground requestAnimationFrame callback delivery, not compositor,
video, or game-rendered frames. `ObservatoryFPS.state()` exposes the SDK version,
status, latest measured window and bounded local history. The graph labels FPS levels and time, and reports the minute’s minimum and largest gap. The window event
`observatory-fps-window` carries exact intervals, start/end, frames, milliseconds,
FPS, p95 and maximum gap; `observatory-fps-state` reports paused/disabled/running.
`ObservatoryFPS.mark("Play")` or `mark("Seek timeline")` records a bounded diagnostic action label without input text. The shared sampler observes long tasks and long animation frames where the browser supports them. A gap of at least 100 ms saves one correlated incident through the existing enrolled browser collector (at most once per ten seconds, one in flight, five-second deadline); unsupported attribution is explicitly recorded. Same-origin script paths have query/fragment removed; no DOM text, inputs, credentials or full page URLs are sent. Apps without a `script[data-project]` enrollment keep evidence local. Off disables sampling, observers and pending collection. Collector completion/failure is visible beside the dip, with a direct retained-trace link on success. These correlations identify recorded work around a dip; they do not prove causality or measure compositor frames.

System health records `fairystack_local_browser_debugging_commands` from FairyStack’s bounded, content-free command receipt health read. This counts accepted Mac browser-attachment recipes in the trailing 24 hours (maximum 1000 receipts), not observed consent dialogs. It remains available independently of browser telemetry coverage.
