Observability
Prometheus metrics, the five Grafana dashboards, 23 alerts, the service level objectives, health checks and logs.
Grounded exposes Prometheus metrics and ships five Grafana dashboards and a set of alerting and recording rules. They work with any Prometheus-compatible stack: Prometheus, Mimir, Cortex or Thanos, with Grafana on top.
| What | Where |
|---|---|
| API metrics | GET /metrics on the internal METRICS_ADDR listener (:9091 in Kubernetes). The public port answers 404. |
| Worker metrics | GET /metrics on WORKER_HTTP_ADDR (:9090). |
| Dashboards | deploy/observability/dashboards/ in the repository (Grafana JSON), or the dashboards component. |
| Alert and recording rules | deploy/observability/alerts/grounded.rules.yaml, or the alerts component. |
| Logs | JSON on stdout, one line per request with route, status, duration_ms and request_id. |
No metric carries a team, user, agent, document or conversation ID. Labels are bounded: route patterns (never raw paths), channels, kinds, outcomes and connection names an admin chose.
Scraping
- Prometheus Operator: add the
monitoringcomponent, a ServiceMonitor for themetricsports ofgrounded-apiandgrounded-worker. Add the label your Prometheus selects on. - Anything else (Prometheus pod discovery, Grafana Alloy, the OpenTelemetry Collector): add
monitoring-annotations.
Both admit the scraper from the monitoring namespace through a NetworkPolicy; patch it to where yours runs. Scrape every 15 to 60 seconds. The dashboards and rules group by the namespace label.
Metrics
The main families (histograms have the usual _bucket, _sum and _count series):
| Area | Metrics |
|---|---|
| HTTP | grounded_http_requests_total, grounded_http_request_duration_seconds by route group, method, route and status; grounded_http_requests_in_flight |
| Chat and retrieval | grounded_chat_answers_total by channel and outcome (ok, no_answer, moderated, model_busy, aborted, error), grounded_chat_first_token_seconds, grounded_chat_duration_seconds, grounded_retrieval_duration_seconds |
| Checks | grounded_systemone_requests_total and its duration by feature; grounded_moderation_decisions_total by stage and decision |
| Model connections | grounded_model_requests_total by connection, kind and outcome (ok, rate_limited, throttled, unavailable, auth, …); grounded_model_request_duration_seconds |
| Ingestion and jobs | grounded_ingest_documents_total by outcome, grounded_crawl_pages_total, grounded_jobs_worked_total and grounded_job_duration_seconds by job kind |
| Install state (workers) | grounded_jobs by kind and state, the oldest waiting and running job's age, grounded_maintenance_mode, open break-glass sessions |
| Governance | grounded_retention_deleted_total, grounded_retention_held, grounded_retention_last_success_timestamp_seconds, break-glass sessions and reads |
| Process | grounded_build_info (version, commit, mode), the Postgres pool, Valkey errors, Go runtime and process metrics |
Metric names may change between minor releases before 1.0; the changelog lists changes.
Dashboards
| Dashboard | Shows |
|---|---|
| Grounded / Overview | The objectives and 30-day error budgets, burn rates, instances up, maintenance mode, the queue, break-glass sessions, firing alerts, chat outcomes, failing model requests and job failures. |
| Grounded / API | Traffic, errors and latency by route group, the slowest and most failing routes, the Postgres pool, Valkey errors, CPU and memory. |
| Grounded / Chat & retrieval | Answers by channel and outcome, time to first token, answer time, retrieval latency, SystemOne latency. |
| Grounded / Ingest & jobs | The job queue, documents by outcome, embedding batch sizes, crawled pages, maintenance mode, retention. |
| Grounded / Models & moderation | Requests, failures, 429s and latency per connection and model kind, SystemOne, moderation decisions. |
Each picks its data source through a DS_PROMETHEUS variable (Mimir works), filters by namespace, and needs Grafana 10.4 or later. Load them by hand (Dashboards → Import), with the Grafana dashboard sidecar (the dashboards component creates one ConfigMap per dashboard labelled grafana_dashboard: "1"; the sidecar must watch Grounded's namespace), with file provisioning, or with the Grafana Operator.
Alerts
The rule file has 23 alerts, each with a severity, a summary, a description and a runbook link:
| Alert | Severity | Fires when |
|---|---|---|
GroundedAPIErrorBudgetBurn / …Slow | critical / warning | 5xx responses burn the 99.9% budget too fast |
GroundedAPILatencyBudgetBurn / …Slow | critical / warning | Requests over 1 s burn the 5% latency budget too fast |
GroundedAPIDown, GroundedWorkerDown | critical | No API or worker scraped for 5 or 10 minutes |
GroundedChatAnswersFailing | warning | Over 10% of answers end in an error for 10 minutes |
GroundedJobsFailing, GroundedJobsDiscarded, GroundedJobStuck, GroundedQueueBacklog | warning | Jobs failing, failing every attempt, running over an hour, or waiting over 15 minutes |
GroundedStateUnreadable | warning | A worker can't read the queue state |
GroundedRetentionFailing, GroundedRetentionStale | warning | A retention kind failed, or hasn't completed in 3 hours |
GroundedMaintenanceModeLong | warning | Maintenance mode on for over 4 hours |
GroundedModelConnectionFailing | critical | Over half of a connection's requests fail for 10 minutes |
GroundedModelConnectionRateLimited | warning | Over 20% refused by 429 or the connection's limit for 30 minutes |
GroundedDBPoolSaturated, GroundedValkeyErrors | warning | The Postgres pool over 90% full; repeated Valkey errors |
GroundedBackupMissing, GroundedBackupJobFailed | critical / warning | The pg_dump CronJob hasn't succeeded for 26 hours; a backup Job failed |
GroundedVolumeFillingUp, GroundedVolumeFull | warning / critical | A Grounded volume over 85% full; over 95% or full within a day |
The backup alerts need kube-state-metrics, and the volume alerts the kubelet's volume stats. A CloudNativePG install backs up through its own scheduled backups; alert on its metrics instead. GroundedAPIDown and GroundedWorkerDown use absent(), so with several installs in one Prometheus, copy them per install with a namespace matcher.
Loading the rules
- Prometheus: list
grounded.rules.yamlunderrule_files. - Prometheus Operator: add the
alertscomponent (a PrometheusRule namedgrounded) and the label yourruleSelectorexpects. - Mimir with Grafana Alloy: the
alertscomponent plus Alloy'smimir.rules.kubernetes, which loads PrometheusRule resources into the Mimir ruler (the CRD must be installed; the Operator isn't needed). - Mimir without Alloy: load the file with
mimirtool rules load.
Service level objectives
| Objective | Measured as | Target |
|---|---|---|
| Availability | API requests (health checks excluded) that don't fail with 5xx | 99.9% over 30 days |
| Latency | Non-streaming API requests that finish within 1 s | 95% |
Streamed chat is outside the latency objective; its speed depends mostly on the model. Uploads and crawls refused during maintenance mode count as 5xx. Your model gateway, identity provider and SMTP relay bound what you can reach, and a single-Postgres install can't meet 99.9% through node or storage failures.
Health checks
| Endpoint | Meaning | Used by |
|---|---|---|
/healthz | The process is running. Never checks dependencies, so an outage doesn't restart pods. | Liveness and startup probes |
/readyz | Postgres, Valkey and object storage answer within 2 s, and the pod isn't draining; otherwise 503 with the failing checks. | Readiness probe |
kubectl -n grounded port-forward svc/grounded-api 8080:80
curl localhost:8080/readyz