Groundeddocs

Observability

Prometheus metrics, the five Grafana dashboards, 23 alerts, the service level objectives, health checks and logs.

Grounded exposes Prometheus metrics and ships five Grafana dashboards and a set of alerting and recording rules. They work with any Prometheus-compatible stack: Prometheus, Mimir, Cortex or Thanos, with Grafana on top.

WhatWhere
API metricsGET /metrics on the internal METRICS_ADDR listener (:9091 in Kubernetes). The public port answers 404.
Worker metricsGET /metrics on WORKER_HTTP_ADDR (:9090).
Dashboardsdeploy/observability/dashboards/ in the repository (Grafana JSON), or the dashboards component.
Alert and recording rulesdeploy/observability/alerts/grounded.rules.yaml, or the alerts component.
LogsJSON on stdout, one line per request with route, status, duration_ms and request_id.

No metric carries a team, user, agent, document or conversation ID. Labels are bounded: route patterns (never raw paths), channels, kinds, outcomes and connection names an admin chose.

Scraping

  • Prometheus Operator: add the monitoring component, a ServiceMonitor for the metrics ports of grounded-api and grounded-worker. Add the label your Prometheus selects on.
  • Anything else (Prometheus pod discovery, Grafana Alloy, the OpenTelemetry Collector): add monitoring-annotations.

Both admit the scraper from the monitoring namespace through a NetworkPolicy; patch it to where yours runs. Scrape every 15 to 60 seconds. The dashboards and rules group by the namespace label.

Metrics

The main families (histograms have the usual _bucket, _sum and _count series):

AreaMetrics
HTTPgrounded_http_requests_total, grounded_http_request_duration_seconds by route group, method, route and status; grounded_http_requests_in_flight
Chat and retrievalgrounded_chat_answers_total by channel and outcome (ok, no_answer, moderated, model_busy, aborted, error), grounded_chat_first_token_seconds, grounded_chat_duration_seconds, grounded_retrieval_duration_seconds
Checksgrounded_systemone_requests_total and its duration by feature; grounded_moderation_decisions_total by stage and decision
Model connectionsgrounded_model_requests_total by connection, kind and outcome (ok, rate_limited, throttled, unavailable, auth, …); grounded_model_request_duration_seconds
Ingestion and jobsgrounded_ingest_documents_total by outcome, grounded_crawl_pages_total, grounded_jobs_worked_total and grounded_job_duration_seconds by job kind
Install state (workers)grounded_jobs by kind and state, the oldest waiting and running job's age, grounded_maintenance_mode, open break-glass sessions
Governancegrounded_retention_deleted_total, grounded_retention_held, grounded_retention_last_success_timestamp_seconds, break-glass sessions and reads
Processgrounded_build_info (version, commit, mode), the Postgres pool, Valkey errors, Go runtime and process metrics

Metric names may change between minor releases before 1.0; the changelog lists changes.

Dashboards

DashboardShows
Grounded / OverviewThe objectives and 30-day error budgets, burn rates, instances up, maintenance mode, the queue, break-glass sessions, firing alerts, chat outcomes, failing model requests and job failures.
Grounded / APITraffic, errors and latency by route group, the slowest and most failing routes, the Postgres pool, Valkey errors, CPU and memory.
Grounded / Chat & retrievalAnswers by channel and outcome, time to first token, answer time, retrieval latency, SystemOne latency.
Grounded / Ingest & jobsThe job queue, documents by outcome, embedding batch sizes, crawled pages, maintenance mode, retention.
Grounded / Models & moderationRequests, failures, 429s and latency per connection and model kind, SystemOne, moderation decisions.

Each picks its data source through a DS_PROMETHEUS variable (Mimir works), filters by namespace, and needs Grafana 10.4 or later. Load them by hand (Dashboards → Import), with the Grafana dashboard sidecar (the dashboards component creates one ConfigMap per dashboard labelled grafana_dashboard: "1"; the sidecar must watch Grounded's namespace), with file provisioning, or with the Grafana Operator.

Alerts

The rule file has 23 alerts, each with a severity, a summary, a description and a runbook link:

AlertSeverityFires when
GroundedAPIErrorBudgetBurn / …Slowcritical / warning5xx responses burn the 99.9% budget too fast
GroundedAPILatencyBudgetBurn / …Slowcritical / warningRequests over 1 s burn the 5% latency budget too fast
GroundedAPIDown, GroundedWorkerDowncriticalNo API or worker scraped for 5 or 10 minutes
GroundedChatAnswersFailingwarningOver 10% of answers end in an error for 10 minutes
GroundedJobsFailing, GroundedJobsDiscarded, GroundedJobStuck, GroundedQueueBacklogwarningJobs failing, failing every attempt, running over an hour, or waiting over 15 minutes
GroundedStateUnreadablewarningA worker can't read the queue state
GroundedRetentionFailing, GroundedRetentionStalewarningA retention kind failed, or hasn't completed in 3 hours
GroundedMaintenanceModeLongwarningMaintenance mode on for over 4 hours
GroundedModelConnectionFailingcriticalOver half of a connection's requests fail for 10 minutes
GroundedModelConnectionRateLimitedwarningOver 20% refused by 429 or the connection's limit for 30 minutes
GroundedDBPoolSaturated, GroundedValkeyErrorswarningThe Postgres pool over 90% full; repeated Valkey errors
GroundedBackupMissing, GroundedBackupJobFailedcritical / warningThe pg_dump CronJob hasn't succeeded for 26 hours; a backup Job failed
GroundedVolumeFillingUp, GroundedVolumeFullwarning / criticalA Grounded volume over 85% full; over 95% or full within a day

The backup alerts need kube-state-metrics, and the volume alerts the kubelet's volume stats. A CloudNativePG install backs up through its own scheduled backups; alert on its metrics instead. GroundedAPIDown and GroundedWorkerDown use absent(), so with several installs in one Prometheus, copy them per install with a namespace matcher.

Loading the rules

  • Prometheus: list grounded.rules.yaml under rule_files.
  • Prometheus Operator: add the alerts component (a PrometheusRule named grounded) and the label your ruleSelector expects.
  • Mimir with Grafana Alloy: the alerts component plus Alloy's mimir.rules.kubernetes, which loads PrometheusRule resources into the Mimir ruler (the CRD must be installed; the Operator isn't needed).
  • Mimir without Alloy: load the file with mimirtool rules load.

Service level objectives

ObjectiveMeasured asTarget
AvailabilityAPI requests (health checks excluded) that don't fail with 5xx99.9% over 30 days
LatencyNon-streaming API requests that finish within 1 s95%

Streamed chat is outside the latency objective; its speed depends mostly on the model. Uploads and crawls refused during maintenance mode count as 5xx. Your model gateway, identity provider and SMTP relay bound what you can reach, and a single-Postgres install can't meet 99.9% through node or storage failures.

Health checks

EndpointMeaningUsed by
/healthzThe process is running. Never checks dependencies, so an outage doesn't restart pods.Liveness and startup probes
/readyzPostgres, Valkey and object storage answer within 2 s, and the pod isn't draining; otherwise 503 with the failing checks.Readiness probe
kubectl -n grounded port-forward svc/grounded-api 8080:80
curl localhost:8080/readyz

On this page