Device fleet anomaly triage
Telemetry anomalies from thousands of devices are clustered, root-cause candidates listed, one incident reaches the right team.
Device fleet anomaly triage
- Alarm: temperature deviation from 312 devices in 20 minutes, region: Marmara warehouses
- Last 6 hours of telemetry pulled; 312 devices split into 3 clusters, largest 284
- Main cluster: same sensor model, same firmware 4.2.1, same gateway group
- Firmware history: 4.2.1 rolled out yesterday 02:00; no deviation on the previous version
- Known issues: 4.2.1 calibration offset bug (KI-0788), workaround: rollback
- Root-cause confidence 0.91: firmware offset; physical cooling failure ruled out
- Incident INC-40917 opened, priority P2, 284 devices attached, root-cause note added
- Bulk rollback command > 100 devices → fleet operations lead approved
- Team summary: cluster, root cause, approved rollback plan and monitoring window
- ■completed
Simulation · derived from real agent definitions · every agent can be built by dialogue with the Autonomous Agent and validated with a test corpus
Anomaly alarm or threshold-breach event from the IoT platform
One incident record, affected device list and recommended action; critical commands approved
Agents
Fleet Health Supervisor
Decomposes the objective, delegates to agents, manages approval points, merges the result.
Signal Agent
Normalises and clusters telemetry
get_telemetrycluster_anomaliesget_device_metadataDiagnosis Agent
Evaluates root cause with firmware, location and event history
get_firmware_historysearch_known_issuescorrelate_eventsOperations Agent
Opens the incident, prepares commands, informs the team
create_incidentprepare_device_commandsend_slackSteps
| # | Kind | Agent | What happens | System | ms | tok |
|---|---|---|---|---|---|---|
| 01 | ingest | Fleet Health Supervisor | Alarm: temperature deviation from 312 devices in 20 minutes, region: Marmara warehouses | — | 200 | — |
| 02 | tool | Signal Agent | Last 6 hours of telemetry pulled; 312 devices split into 3 clusters, largest 284 | Time-series DB | 520 | — |
| 03 | reason | Signal Agent | Main cluster: same sensor model, same firmware 4.2.1, same gateway group | — | 640 | 700 |
| 04 | tool | Diagnosis Agent | Firmware history: 4.2.1 rolled out yesterday 02:00; no deviation on the previous version | IoT Platform (MQTT / Digital Twin) | 440 | — |
| 05 | retrieve | Diagnosis Agent | Known issues: 4.2.1 calibration offset bug (KI-0788), workaround: rollback | Known Issues (vector) | 480 | 520 |
| 06 | verify | Diagnosis Agent | Root-cause confidence 0.91: firmware offset; physical cooling failure ruled out | — | 560 | 480 |
| 07 | write | Operations Agent | Incident INC-40917 opened, priority P2, 284 devices attached, root-cause note added | ITSM / Incident | 480 | — |
| 08 | approval | Fleet Health Supervisor | Bulk rollback command > 100 devices → fleet operations lead approved | — | 2,600 | — |
| 09 | notify | Operations Agent | Team summary: cluster, root cause, approved rollback plan and monitoring window | Slack | 240 | — |
Staged OTA update rollout
Firmware rolls out in 1% → 10% → 100% stages; a health gate at every stage, human approval for critical fleets.
Cold-chain excursion response
A temperature excursion is assessed instantly, product risk computed, a field crew dispatched, the quality decision recorded.
Time to move from experimenting with AI to transforming with it.
In a 30-minute discovery session we take your 2–3 priority business problems, show a live demo of a similar scenario, and draft a roadmap that starts with the Value layer.
- Your 2–3 priority problems
- Live demo of a similar scenario
- Roadmap starting with value analysis