build observability

Watch every remote build in real time.

A live view over your Buildbarn RBE cluster — cache hit rate, worker saturation, queue depth, execution latency — read straight from the Prometheus it already exposes. No new database, no dashboard to hand-build, and the same numbers are agent-callable tools.

available
Cache / workers / queue / latency, liveReads the Prometheus your RBE already exposesEvery metric is an agent-callable tool tooZero setup — no dashboard to build or maintainNo Prometheus? Panels go empty — never a fake number
0
extra infra to run
reads your existing Prometheus
4
live dashboards
overview · cache · workers · queue
5
metrics as agent tools
the same numbers, callable
32Mi
memory request
10m CPU · one replica
why it's different

Your cluster already emits everything. Read it.

Every Buildbarn pod exposes Prometheus metrics at :9980/metrics, and you're already scraping them. Standing up a bespoke dashboard to see them is busywork.

builds points at your Prometheus (BUILDS_PROMETHEUS_URL) and runs the PromQL — instant and range queries over a rolling one-hour window — that turns raw Buildbarn counters into the four views you actually watch during a build: cache hit rate, CAS/AC throughput, worker saturation by platform, and scheduler queue depth with execution-latency percentiles. It holds no state of its own. And because every panel is backed by a typed query, the same numbers are exposed as MCP tools — so an agent can check the cache hit rate or the worker count exactly the way the console does.

It doesn't scrape a new thing. It reads the Prometheus your Buildbarn cluster already exposes.

plugin-builds · reads the cluster you already run
the console

The whole cluster, at a glance.

The Overview panel is a KPI strip: one tile per signal, each with its hour-ago delta (the last sample vs the first in the window). Values below are illustrative — the panel renders whatever your cluster reports.

Overview live
/api/gw/builds/overview · rolling 1h window · 60s step
signalvaluevs 1h ago
cache hit rate 87% ▲ 4 pts
active workers 3 steady
queue backlog 0 drained
median exec latency 340ms ▼ 60ms
artifacts served (1h) 12,481 —
bytes served (1h) 1.2 GiB —
source: Buildbarn Prometheusrefresh: on panel load, fresh query
the cache

Every read and write, by store.

The Cache panel breaks throughput out by store — Content-Addressed Storage and the Action Cache — and tracks the action-cache hit rate that decides whether a build is fast or cold. Series and rates below are illustrative.

Cache
CAS · AC · FSAC · rate over [5m]
seriesvalueprometheus source
AC hit rate (GetActionResult) 87% grpc_server_handled_total
CAS read 420 ops/s blob_access_operations_blob_size_bytes_count
CAS write 96 ops/s blob_access_operations_blob_size_bytes_count
AC read 210 ops/s blob_access_operations_blob_size_bytes_count
AC write 44 ops/s blob_access_operations_blob_size_bytes_count
hit-rate domain: 0 → 1fallback: AC reuse when GetActionResult isn't scraped
the pipeline

Scrape → query → shape → render.

Four stages, no state in the middle.

01

Scrape

Every Buildbarn pod (scheduler, workers, storage) exposes Prometheus at :9980/metrics. Your Prometheus already scrapes it — builds adds nothing to the collection path.

02

Query

On each panel load, builds runs PromQL instant and range queries against BUILDS_PROMETHEUS_URL — rates over [5m], a rolling 1h window at a 60s step (60 points). No new time-series database.

03

Shape

Results are flattened into typed builds.v1 DTOs and Vega samples — the same shape whether the caller is a panel or an MCP tool.

04

Render

Vega-Lite meridian panels draw the charts in the console; the identical numbers answer the overview, cache_hit_rate, list_workers, cache_stats, and list_invocations tools for agents.

grounded, not guessed

Every panel names its metric.

These are the real Buildbarn series each view is computed from — verified against a live cluster's :9980/metrics, because the earlier metric names most dashboards copy don't exist in current builds.

signal → metric
buildbarn · InMemoryBuildQueue · blobstore
panelprometheus metricwindow
cache hit rate grpc_server_handled_total{grpc_method=GetActionResult} rate [5m]
cache throughput ..._blob_access_operations_blob_size_bytes_count rate [5m]
workers by platform ..._build_queue_workers_created_total − ..._removed_total sum by(platform)
queue backlog ..._tasks_scheduled_total − ..._tasks_completed_..._count instant
exec latency p50/p90/p99 ..._tasks_executing_duration_seconds_bucket histogram_quantile
verified: against a live :9980/metrics
vs a DIY dashboard

Not another dashboard to maintain.

The alternative is standing up and babysitting a dashboards deployment, guessing the metric names, and leaving agents out entirely.

fastverk builds
a hand-built dashboard
Setup
✓ Point at your Prometheus URL
Build + version dashboards by hand
Metric names
✓ Verified against the live :9980/metrics
Guess, then fix when panels are blank
Agent access
✓ Same metrics as 5 MCP tools
Humans only, in a browser tab
Footprint
✓ 10m CPU · 32Mi · one replica
A dashboard stack to run
Where it lives
✓ A native panel in the same console
A separate tool and login
No data
✓ Empty panels, never an error
Broken panels or stale caches

No Prometheus reachable? Every panel degrades to an empty state — never an error, never a fabricated number.

plugin-builds · degrade to empty

source · github.com/fastverk/plugin-builds · tools: overview, cache_hit_rate, list_workers, cache_stats, list_invocations

Prove it's safe to merge.

builds is one of 14 plugins in the fastverk console — hosted, or in your own cloud.