Case study 01
Telecom & IoT
Monitoring 50,000+ devices across the country
Our engineers scaled a high-availability monitoring platform that watches 600,000 items (metrics) in real time at more than 100 sites nationwide.
- devices monitored
- 50,000+
- sites nationwide
- 100+
- monitoring stack, no single point of failure
- HA
- Zabbix
- High availability
- Distributed proxies
- Custom dashboards
The challenge
The client ran a large fleet of devices spread across the country. Problems were often found late, and many needed a field team on site just to diagnose them. The existing monitoring setup could not keep up with the number of devices.
What we built
- Redesigned the Zabbix architecture for high availability and horizontal scale
- Split collection across proxies close to each region to cut load and latency
- Built custom dashboards for the network operations center (NOC)
- Designed alert routing so the right team gets the right alert, with less noise
Results
- Incidents detected in minutes instead of being reported by users
- Fewer site visits: most issues are now diagnosed remotely
- One monitoring platform that scales with the device fleet
- Clear NOC dashboards with a nationwide view at a glance
