Observability#
Monarch provides three complementary job-local observability surfaces: distributed telemetry for SQL analysis, the Monarch Dashboard for visual monitoring, and the Mesh Admin TUI for live diagnostics. Configure telemetry and mesh admin together to connect all three data paths. OpenTelemetry provides a separate export path for external backends.
Choose a surface#
Use case |
Surface |
What it provides |
|---|---|---|
Query actor, mesh, message, and trace history |
DataFusion SQL across host-local collectors |
|
Monitor topology, status, failures, and traffic in a browser |
Live summaries, hierarchy views, and a job DAG |
|
Diagnose a running mesh from a terminal |
Live topology, health checks, and py-spy stack traces |
|
Export metrics and logs to an external backend |
OTLP export to systems such as Prometheus and Loki |
The dashboard reads distributed telemetry. The TUI reads the Mesh Admin API. Periodic mesh-admin snapshots also flow into telemetry so the dashboard can show the live administrative topology.
Distributed telemetry is a job-scoped, in-memory query system. OpenTelemetry export is the external aggregation path for metrics and logs. You can use both on the same job.
Enable observability#
Configure observability before calling state():
from monarch.job import ProcessJob, TelemetryConfig
job = (
ProcessJob({"workers": 2})
.enable_admin()
.enable_telemetry(
TelemetryConfig(
retention_secs=600,
include_dashboard=True,
dashboard_port=8265,
snapshot_interval_secs=30,
)
)
)
state = job.state(cached_path=None)
enable_telemetry() starts distributed telemetry and the optional dashboard;
enable_admin() starts the API used by the TUI. Configuring both also enables
periodic introspection snapshots when snapshot_interval_secs is positive.
Adapt the example to your existing LocalJob, ProcessJob, SlurmJob, or
KubernetesJob rather than creating a second job.
The returned state exposes each surface:
print(state.dashboard_url)
print(state.admin_url)
client = state.query_engine_client
if client is None:
raise RuntimeError("telemetry did not start")
result = client.query(
"SELECT new_status, COUNT(*) AS count "
"FROM actor_status_events GROUP BY new_status"
)
print(result["rows"])
Use state.dashboard_url in a browser. Pass state.admin_url to
monarch-tui --addr.
Data flow#
actor and proc events ──> host-local collectors ──> distributed query engine
├─> SQL query client
└─> Monarch Dashboard
live mesh ──> Mesh Admin API ──────────────────────────────> Mesh Admin TUI
└──────> periodic snapshots ──> distributed telemetry ─> Dashboard DAG
Telemetry collection and distributed query fan-out are best-effort. A failed or slow worker collector can reduce query coverage without failing the job. Telemetry tables are in memory and belong to the current job allocation; they are not a durable event archive.
Configuration#
|
Default |
Effect |
|---|---|---|
|
|
Retention window for message tables; |
|
|
Advertise the browser dashboard |
|
|
Preferred dashboard port; use |
|
|
Mesh-introspection snapshot interval; |
For an admin server without distributed telemetry, call job.enable_admin().
That supports the TUI, but it does not provide the dashboard or SQL history.