Time aggregation determines how multiple data points within a collection/aggregation interval are combined before evaluation.
Common Time Aggregation Methods
| Aggregation | Use Case | Example |
|---|---|---|
| Latest | Point-in-time value of a gauge | Last reported heartbeat timestamp |
| Max | Peak values, worst-case scenarios | Container restarts, max CPU spike |
| Min | Minimum thresholds, availability | Minimum available memory |
| Avg | General trends, smoothed metrics | Average response time |
| Sum | Total counts, cumulative metrics | Total requests, error count |
| Count | Number of occurrences | Event frequency |
| Count Distinct | Unique values | Unique users, distinct IPs |
| P50/P95/P99 | Latency percentiles | Response time distributions |
| Rate | Changes per time unit | Requests per second, errors per minute |
| Increase | Absolute growth over time period | Total value growth since previous measurement |
To learn more about aggregations in metrics, visit the Metric types and aggregation.
Aggregation Examples
For Container Restarts:
Metric: k8s.container.restarts
Aggregation: max (use running_diff to compute increments)
Evaluation: at least once
Formula: running_diff(k8s.container.restarts, cutoff_min=0)
Reason: Restart counters are cumulative and don't reset to zero immediately after a restart, which can make alerts continuously fire.
Compute the difference between consecutive samples and drop negative values (caused by counter resets) using cutoff_min=0.For Memory Usage:
Metric: system.memory.usage
Aggregation: avg or max (depending on use case)
Formula: (used - cached) / total * 100
Reason: Exclude cached memory for accurate usage representationFor Latency Monitoring:
Metric: http.server.duration
Aggregation: P95 or P99
Evaluation: on average (not in total)
Reason: Percentiles shouldn't be summed; average P95 over time makes senseFor Throughput Analysis:
Metric: http.requests.total
Aggregation: sum
Evaluation: in total
Reason: Provides the total number of requests over the evaluation window, useful to monitor system loadFor Error Rate Calculation:
Metric: system.errors
Aggregation: count or rate
Evaluation: at least once
Reason: Detects any occurrence of error spikes and evaluates system reliabilityDetecting Trends with Running Diff
You can detect continuous uptrends in derived metrics by using the running_diff function on individual queries rather than on formulas. This is useful for scenarios like monitoring RabbitMQ message queue trends where you want to alert on sustained increases rather than temporary spikes.
Setting Up Uptrend Detection
To detect a continuous uptrend in a derived metric (like the rate difference between messages published and acknowledged), follow these steps:
-
Create your base queries:
- Query A:
rate(messages_published)(every 120s) - Query B:
rate(messages_acked)(every 120s)
- Query A:
-
Apply running difference to each query:
- Apply
running_diffdirectly on Query A with 120-second interval - Apply
running_diffdirectly on Query B with 120-second interval
- Apply
-
Create the formula:
- Use formula:
A - Bto get your derived rate difference
- Use formula:
-
Set up the alert condition:
- Alert when the result is
> 0to detect upward trends
- Alert when the result is
The running_diff function calculates R(t) - R(t-previous), allowing you to detect when each consecutive reading is higher than the previous one, which is exactly what you need for sustained uptrend detection.

Detecting Metric Stagnation
Use running_diff for the inverse case too, when a metric that must keep increasing stops changing. A heartbeat metric that reports a Unix timestamp works this way, because each reading carries a later timestamp than the reading before it.
-
Create the query:
- Select your gauge metric, for example
app.heartbeat.timestamp. - Set time aggregation to Latest, which suits a gauge that reports a point-in-time value.
- Filter the query to one source, or group by the attribute that identifies the source, such as
host.name. - Add the
running_difffunction.
- Select your gauge metric, for example
-
Set up the alert condition:
- Alert when the result is Equal to
0, with match type All the Time over your chosen rolling window.
- Alert when the result is Equal to
Space aggregation combines every matching series into one value. If you skip the filter and the group-by, the sources that still report keep that combined value moving, and you never see the frozen one.
When every running_diff value in the window equals 0, no consecutive pair of readings changed.
running_diff subtracts the previous point from each point, and it drops the first point, so it returns one value fewer than the points it receives. For a query that uses running_diff, SigNoz reads one step before the start of your window, so the earlier point can come from that step. Make sure that the window and that extra step together hold two aggregated points. Otherwise the query returns no value, and the alert cannot fire.