Monitoring & alerts
Monitoring shows how your servers are doing and tells you in Telegram or by email when something needs attention. It lives in the Cloud Panel under Monitoring and costs nothing extra.
This page covers where the numbers come from, how alert rules work, how an alert behaves from start to resolve, and how disk auto-extend grows a root disk before it fills up.
Where the data comes from
There are two sources. The hypervisor sees your server from outside, the guest agent reports from inside.
| Metric | Source | How often |
|---|---|---|
| CPU, disk I/O, network | Hypervisor | Every few seconds |
| Disk space per mount, memory, load average | Guest agent | Once a minute |
| Disk latency, CPU split (user, system, iowait, steal) | Guest agent | Once a minute |
| Memory and I/O pressure (PSI), OOM kills | Guest agent | Once a minute |
Hypervisor metrics work on every server. Guest metrics need the agent running and the Monitoring switch on in the server's Overview, under Agent & monitoring. Both are on by default for servers from our OS templates. A server installed from your own ISO has hypervisor metrics only.
Pressure metrics need a kernel that reports PSI. Where it does not, those charts stay empty and pressure alerts never fire.
The Monitoring page
- Overview. Charts across all monitored servers.
- Server status. One row per server with its current CPU, memory, disk space and agent state.
- Alerts. Active alerts on top, then history, with search, filters and sorting.
- Alert settings. Rules, silences and where alerts go.
A server with Monitoring switched off is not on this page and has no alerts. Its hypervisor charts stay on the server's own page.
When there is an active alert that is not silenced, a red mark appears next to Monitoring in the sidebar and on the Alerts tab.
What can be alerted on
| Alert | What it measures | Optimized rule |
|---|---|---|
| High CPU | Average CPU of the server over the rule's time | Warning 80%, critical 90%, 5 min |
| High memory | Memory in use without page cache | Warning 90%, critical 95%, 5 min |
| Disk space | Used space of every mount, /boot excluded. One alert per mount | Warning 80%, critical 90% |
| Disk latency | Average I/O latency of each disk inside the server | Off. Warning 50 ms, critical 200 ms, 5 min |
| Load average | One-minute load divided by the number of vCPUs | Off. Warning 1.5, critical 3, 5 min |
| Memory pressure | Share of time processes wait for memory (PSI some) | Off. Warning 20%, critical 50%, 5 min |
| IO pressure | Share of time processes wait for disk I/O (PSI some) | Off. Warning 20%, critical 50%, 5 min |
| OOM kill | The kernel killed a process for lack of memory | On |
| Guest agent not responding | The agent stopped answering on a running server where it worked before | On, after 3 min |
| CPU Guard | Burst servers only. CPU Guard lowered the CPU limit, and later let go | On |
The time on a threshold rule is how long the reading is averaged: 1, 3, 5, 10, 15 or 30 minutes. Disk space has no time and fires on the latest reading.
Alert rules
A rule is an alert kind, a warning and a critical threshold, a time, a scope and the channels it sends to. A new account has no rules.
The scope decides which servers a rule covers:
- All servers, including servers you create later.
- A tag, written as
key=value, or justkeyfor any value. - Selected servers.
Any scope can exclude tags. For example High CPU for all servers except role=batch, plus a stricter High CPU for env=prod.
Several rules of the same kind can cover one server, and we check all of them. The server still gets one alert per kind (and per mount or disk) with the severity of the strictest rule that fired. The message lists every rule that fired.
Optimized rules
We made our own set of rules that fits most servers, and we recommend starting with it. The thresholds are in the table above: they catch real trouble and stay quiet on normal load. Change them later if your servers need something else.
Load optimized rules in Alert settings creates this set as ordinary rules for all servers, which you can then edit or delete. It replaces all rules the account has. Delete all rules removes them.
When an account creates its first server with Monitoring on and has no rules yet, we load the optimized set for it. That happens once per account. If you delete the rules later, they do not come back.
Setting up your own rules
Open Monitoring, Alert settings, and press New rule (Add rule manually when the account has none). The fields:
| Field | What it does |
|---|---|
| Name | Optional. Shown in alert messages, handy when several rules of one kind cover a server |
| Alert | The alert kind from the table above |
| Servers | All servers (new ones included), servers with a tag, or selected servers |
| Except servers with tags | Servers with any of these tags are left out of the rule |
| Warning at, Critical at | The two levels. Leave one empty to skip that level |
| Average over | How long the reading is averaged before it is compared with the levels |
| Send to | Telegram, email or both |
Example 1. Stricter CPU for production. The optimized High CPU rule covers all servers at 80% and 90%. You want to hear earlier about servers tagged env=prod.
- Alert: High CPU
- Servers: servers with a tag,
env=prod - Warning at 70, Critical at 85, Average over 5 min
- Send to: Telegram
Both rules cover a production server. At 75% the new rule fires a warning. At 85% it goes critical, while the optimized rule would wait until 90%. You still get one alert per server that names both rules when both fired.
Example 2. No CPU alerts for batch servers. Servers tagged role=batch run at 100% CPU on purpose. Open the optimized High CPU rule and add role=batch to Except servers with tags. Every other server keeps the rule, the batch servers never get a High CPU alert. Other alerts, like Disk space, still cover them.
Example 3. Database disk, critical only, to everyone. For the database server you want one loud alert and no warnings.
- Name:
db disk - Alert: Disk space
- Servers: selected servers, your database server
- Warning at empty, Critical at 85
- Send to: Telegram and email
The optimized Disk space rule still sends its warning at 80% to its own channel. At 85% this rule makes the alert critical and it also goes to the owner's email. If you want this server on your own rule alone, add a tag such as role=db to it and exclude that tag in the optimized rule, like in example 2.
How an alert behaves
An alert sends one message when the problem starts, one more if it goes from warning to critical, and one when it is resolved. It does not repeat while the problem lasts.
- Resolve with a margin. An alert resolves when the reading drops 5 points below the warning threshold of every rule that covers the server (0.5 for load average), so a server hovering at the threshold does not send a message every minute.
- Coming back soon. If the same alert opens again within 30 minutes after it resolved, it continues quietly. Disk space, OOM kill and CPU Guard are the exceptions, for them a repeat is news.
- Stopped or deleted servers. Open alerts of a server that stopped or was deleted resolve without a message. That is not a recovery.
- No agent data. If agent data stops coming from a running server, alerts based on it stay as they are until data returns. The Guest agent alert covers that case.
- OOM kill resolves quietly 15 minutes after the last kill.
- CPU Guard. While the guard holds a burst server, its CPU chart shows the limit and not the real demand, so High CPU does not open or resolve during that time. A new hold within an hour after the release continues the old alert without a message.
A message carries the server name, its public IP (floating, or private when there is none), the reading, the thresholds and the rules that fired, with a link to the server.
Where alerts go
Each rule sends to Telegram, email or both.
- Telegram goes to everyone in the account who linked @serverscamp_bot and has Servers notifications on in the bot.
- Email goes to the account owner.
Optimized rules send to Telegram when someone in the account has it linked, and to the owner's email otherwise. Alert settings shows who gets what, with a Send test button for each channel.
One Telegram account links to one ServersCamp account. To move it to another account, send /unlink to the bot first. When the last working Telegram leaves an account (unlinked, the bot blocked, or Servers switched off), rules that sent to Telegram switch to the owner's email, so alerts do not go quiet. When a Telegram works there again, they switch back by themselves. Rules you set channels on by hand stay as you set them.
Silence
A Telegram message about a problem has Silence 1 h and Silence 24 h buttons. They mute that alert kind on that server. Silences can also be set and removed in Alert settings. A silenced alert is still recorded and shown on the Alerts tab, it only sends nothing.
Message limits
Each account gets at most 20 Telegram messages an hour (100 a day) and 6 emails an hour (30 a day). When a limit is reached, you get one message saying alerts are paused and until when. Every alert is still on the Monitoring page.
Disk auto-extend
Auto-extend grows the root disk, partition and filesystem when / reaches the threshold you set. It uses the same guest agent data as the Disk space alert and needs Monitoring on. When both are set up, the Disk space alert's Resolved message says the disk grew, so you get one message instead of two. Settings, limits and what stops it are in Disk auto-extend.