MONITORING

The watch never blinks.

Health checks, metrics, alerts and live log streaming across every resource Talos manages — with a dashboard shaped for what you are actually running rather than one generic view for everything.

Install itSee the install guide
SSE·Health checks·Per-type dashboards·Email / webhook alerts

THE MANUAL WAY

Finding out from the customer.

THE SILENT FAILURE

A nightly job has failed for eleven days. It exits non-zero into a log nobody reads, and the first sign of trouble is a customer email.

THE WRONG DASHBOARD

Every project shows the same four charts, none of which are the numbers that matter for a database, a cluster or an ERP site.

THE ALERT NOBODY TRUSTS

Alerts fire so often that the channel is muted. The one that mattered was in there somewhere.

WHAT TALOS DOES

Signals worth reacting to.

Per-type dashboards

A database project shows backup age and verification status; a cluster shows node pressure. Each project type has its own view rather than a shared default.

Health checks

Reachability, resource pressure and service state, polled on a schedule and recorded so trends are visible.

Live log streaming

Follow any running job or container as it happens. The stream is the same thing that lands in the record.

Alerts that name the resource

Alerts are scoped to a resource and an organization, so they reach the people responsible for it and nobody else.

Job failure attribution

A failed job points at the change, the parameters and the person, rather than at a stack trace alone.

Fleet overview

One page for the whole estate, with the exceptions surfaced instead of buried in a list of green.

talos.internal/monitoring
Servers
Kubernetes
Frappe
Docker
Databases
Pipelines
Monitoring
The Watch86 resources · 2 alerts
acme.erpHEALTHY
cpu 18% · mem 41% · disk 62%
prod-eu · nodesHEALTHY
12 ready · 0 pressure
legacy-crm backupALERT
unverified for 6 days
db-01 · diskALERT
88% used · +4%/week
! backup verification failed · legacy-crm
→ last verified 2026-07-28
→ notified 2 people with db.backup access
! db-01 disk 88% · projected full in 21d
✓ 84 resources healthy

HOW IT RUNS

Every action is a tracked job.

Requestyou, API or schedule
Queuegated if required
the Sentinelbackground worker
AdapterSSH · Ansible · Terraform · API
Your infrastructurelogs stream back live

WHAT MAKES THIS DIFFERENT

An alert that reaches the wrong organization is worse than none.

Alerts used to be the one part of the system with no tenant boundary — which meant an alert raised for one client’s infrastructure could surface in another client’s console. That is fixed, and it is worth stating plainly rather than quietly: alerts are scoped to the organization that owns the resource, and the scoping is enforced in the data path rather than by a filter each view has to remember.

Alerts route by who holds permission on the resource, not by a global notification list.
Every project type has a dashboard built for it — eight of the ten previously shared one that suited none of them.
Clusters attached by kubeconfig are monitored like any other, rather than silently skipped.
alert routing · legacy-crm
resource scoped
organization: northwind-ops
OK
recipients resolved
2 people hold db.backup on this resource
SENT
other organizations
not notified · not visible
SCOPED
recorded
alert and its resolution in the ledger
DONE

RELATED CAPABILITIES

Servers

The hosts most of these signals come from.

explore →

Databases

Backup age, drift and verification failures.

explore →

AI Engine

Turn a failing job into a reviewable proposed fix.

explore →

Watch one thing properly.

Connect a server and the health checks start immediately. No agent, no exporter to configure.

Install in one commandSee every capability →