Metrics, logs, traces and alerts live in separate systems, so investigations mean switching tools and losing context. Duplicate alerts bury the events that matter and wear out the on-call rota. Root-cause analysis depends on individual experts, keeping MTTR high. Routine inspection is done by hand, with inconsistent scope and conclusions that are never kept.
SupPulse connects metrics, logs, traces, alerts and change context and lets you query, drill down and correlate them from one console. AI-assisted diagnosis calls the observability tools on demand, collects evidence step by step and forms a root-cause hypothesis; scheduled inspection checks the connected infrastructure on a plan, produces a report you can revisit and pushes it to the team's messaging channel. Existing data sources and collectors stay in place; no observability store has to be migrated.
Capabilities
- AI-assisted diagnosis and Q&A: start an investigation in natural language; an agent calls the observability tools, gathers evidence and forms a root-cause hypothesis. Multiple models, expert skills and MCP integrations are configurable; session history and an agent audit trail are kept
- Unified observability: query metrics (PromQL), logs (Elasticsearch / Loki), traces and dashboards from one console; APM shows service topology and RED metrics, with drill-down to traces and spans
- Scheduled infrastructure inspection: plans by cron and time zone, custom scope, run-now; reports are retained and pushed to Feishu, DingTalk, WeCom or a webhook
- Anomaly detection and alert correlation: dynamic baselines and alert rules flag anomalies; rule labels link upstream and downstream alerts, and a wait window suppresses downstream notifications; observe-only, enforce and emergency-release modes, with every decision auditable
- Event centre and operations graph: ingest change and runtime events from GitLab, Jenkins, ArgoCD, Kubernetes and more, with search and a timeline; with the optional graph (Neo4j) enabled, the AI can query service dependencies and recent changes through read-only tools
- Edge collection and platform governance: OpenTelemetry Collector and OpAMP manage collection nodes; plugins cover hosts, middleware, network and databases; multi-tenancy, IAM, audit and configuration versioning are built in