Aggregation types
Conversion, count, sum, mean, ratio — pick the right aggregation, set outlier handling, and avoid the variance traps that bite means and sums.
A metric is "events aggregated per user." The aggregation function turns each user's event stream into a single number the platform charts over time and alert rules compare. Pick the right one and your power calculations are honest; pick the wrong one and a single outlier can swing your result.
The five types
| Type | Per-user value | Best for | Variance behaviour |
|---|---|---|---|
| rate | 1 if the numerator happened, else 0 | Did they buy? Did they retain? | Bounded p(1-p) — friendliest. |
count | Number of events | Page views, sessions, clicks | Long-tailed; outliers possible. |
sum | Sum of a numeric property | Revenue, time spent, items added | Heavy-tailed; outliers a real problem. |
mean | Average of a numeric property | Order value, session length | Same as sum, plus zero-handling. |
p | A percentile of a numeric property | Latency, payload size | Charted exactly; averaged per user. |
Conversion — write it as a rate
The simplest and statistically the friendliest, and it is a division: how many did the thing, out of how many could have.
shipeasy metrics create purchase_conversion \
--event-name purchase \
--query 'count(purchase) / count(session_start)'On a chart that is the rate over each bucket. In an experiment it collapses per user to a 0 or
a 1 — did the numerator happen, among the users eligible for the denominator — so the metric
mean is the fraction who converted, variance is bounded by p(1-p) capping at 0.25, power
calculations are cheap and no single user can be an outlier.
Use it whenever the question is yes/no.
There is no single-event conversion aggregation. count_users(purchase) used to be one and was
removed: Analytics Engine samples rows under load and weights the survivors, and a distinct count
cannot be reweighted — sampling reweights rows, not sets — so it under-reported by the sample
interval with no way to correct it. A rate has no such problem, because both sides are sampled the
same way. Metrics stored before the removal still chart; new ones are written as rates.
Count
Per user: how many of the event happened?
shipeasy metrics create sessions_per_user \
--event-name session_start \
--query 'count(session_start)'Each user contributes an integer ≥ 0. Means and variances behave reasonably for low-volume events (0–5 per user) and get long-tailed for high-volume ones (a power user with 200 sessions skews the mean). For long-tailed counts, consider:
- Tightening the winsorize percentile (
--winsorize 95instead of the 99 default) to clamp the heavy tail. - Switching to a rate (
count(session_start) / count(app_open)— did they start a session at all?) if the count distinction doesn't drive your decision.
Sum
Per user: sum a numeric property across all matching events.
shipeasy metrics create revenue_per_user \
--event-name purchase \
--query 'sum(purchase, revenueCents)'The classic "did this change make us more money per user." Each
user contributes their total revenue. Non-purchasers contribute 0.
Sums are heavy-tailed. One $50,000 enterprise purchase can swing
the mean for thousands of users. The metric's --winsorize percentile
clamps that — the CLI default is --winsorize 99 and is rarely wrong.
Mean
Per user: average a numeric property across their events.
shipeasy metrics create avg_order_value \
--event-name purchase \
--query 'avg(purchase, revenueCents)'There's a trap with avg: should users with zero matching events count
as 0, or be dropped from the denominator entirely? Today the DSL's
avg averages across all exposed users — non-purchasers contribute 0
("average revenue per exposed user"). To answer the "average across
buyers only" question instead, define two metrics and divide via the
ratio aggregation, or compute the per-buyer average offline.
Rate
Per user: numerator-event count divided by denominator-event count. Division is
ordinary — write a / b. (ratio(a, b) is the same thing spelled as a
function, and still parses.)
shipeasy metrics create click_through_rate \
--event-name click \
--query 'count(click) / count(impression)'As an experiment metric both sides must be count: the per-user collapse
turns the rate into "did the numerator happen, among the denominator-eligible
users", and that question only has an answer when both sides count events. On a
chart there is no such restriction — sum(order, amount) / count(session)
draws fine.
Use ratios for inherently-ratio questions: clicks per impression, conversions per visit, errors
per request. Don't compute the ratio yourself and store it as a mean — the math is
different.
Why it matters: the naive ratio-of-means (mean of numerators divided by mean of denominators)
under-states the variance. Shipeasy uses the delta method to compute it correctly — which is
what makes the confidence band the dashboard draws, and any anomaly rule armed on the series,
trustworthy rather than jumpy. If you had computed clicks/impressions yourself per user, then
taken the mean, you'd be calculating the wrong thing.
The dashboard shows numerator and denominator alongside the ratio so the math is auditable.
Outlier handling
Sums and means need outlier handling. Two options:
Winsorize at a percentile. Anything above is clamped to the value at that percentile.
shipeasy metrics create revenue_per_user \
--event-name purchase \
--query 'sum(purchase, revenueCents)' \
--winsorize 9999 (default), 95, 90. Lower percentiles clamp more aggressively.
Trims the rarest 1% (or 5% / 10%) to the value of that percentile,
keeps the body of the distribution intact.
For a hard ceiling (e.g. session length can't reasonably exceed 4 hours, revenue per user can't exceed your enterprise plan price.
The cutoff is computed from the combined control + treatment sample, then applied to both arms. This prevents the bug where one variant accidentally has its outliers preserved and looks artificially better. You can verify in the dashboard: the "applied threshold" row shows the same value across arms.
The trade-off: clamping reduces variance (good — narrower CI, easier to detect lift) at the cost of slightly understating the true effect when the variant genuinely moves the tail (rare).
Filters
You can tighten what counts toward a metric with filters. Same shape as feature flag targeting rules,
applied to the event's properties payload before aggregation:
shipeasy metrics create organic_purchase \
--event-name purchase \
--query 'count(purchase{channel="organic"})'Now only purchase events with channel == "organic" count. Multiple
predicates inside {} are ANDed (commas separate them, and or groups them).
String values must be double-quoted in the DSL. =~ matches a glob, not a
regex: * is any run of characters, ? is exactly one, and the pattern matches
the whole value.
Common shapes:
# Web only (exclude mobile app purchases)
'count(purchase{platform="web"})'
# Geographic slice — a value set, not a regex
'count(purchase{country in ("US", "CA", "GB")})'
# Everything EXCEPT free tier — one negation, in front of the predicate
'count(purchase{not tier="free"})'
# Multiple — ANDed inside the same {}
'count(purchase{platform="web", country="US"})'Filters run cheaply during aggregation. You can have many metrics that share the same underlying event, each filtering differently — no need to log the same purchase twice with different names.
Direction
The dashboard infers each metric's "good" direction from its name and its
aggregation type. Words like error, fail, latency, crash, churn,
bounce are read as lower-is-better; that decides which way a trend cell is
coloured and which side a directional alert rule watches. Set direction
explicitly on the metric when the heuristic gets it wrong.
When you don't have the event yet
A metric definition can be created before any matching events exist. The pipeline picks them up
once they start arriving, so you can define the metric for an upcoming change ahead of time and
deploy the track() call as part of the same release.
The series only ever covers events that were actually logged — creating the metric does not back-fill a period during which nothing emitted the event.
Inspecting a metric
# What does the metric definition look like?
shipeasy metrics show purchase_conversion
# What's the historical baseline?
shipeasy metrics series purchase_conversion --from 2026-06-01 --to 2026-06-30 --bucket dayRun shipeasy metrics series before you put a rule on a metric — the daily buckets show you the
baseline rate and how much it already wanders, which is what a sane threshold has to sit outside
of.