Skip to main content
Question

Tuesday's Tip of the Week - Dashboards for Detection Volume, Rule Health, and Data Quality

  • October 6, 2026
  • 0 replies
  • 3 views

dnehoda
Staff
Forum|alt.badge.img+19

October 6, 2026 

 

Audience: SOC leads, detection engineers, SecOps platform admins


Summary: Google SecOps native dashboards turn detections, rules, and ingestion telemetry into trends, so you can tell whether your detection program is working, degrading, or broken. This briefing defines three core dashboards, the metrics behind each, and the YARA-L 2.0 queries to build them. All queries are based on Google's published dashboard examples.

1. Purpose

Triaging alerts one at a time doesn't show problems that build up across the whole program: telemetry that has stopped flowing, parser regressions, rules that have drifted, and growing alert fatigue.

Question Signal
Is the program working? Alert volume and severity mix stay inside baselines, and high-value rules keep firing
Is it degrading? Noise is growing, precision is falling, parsing errors are rising
Is it broken? A log source has gone silent, rules never fire, or a parser is failing

 

Data sources used

Data source Query prefix Lookback Used for
Detections detection. 365 days Volume by severity, top rules
Rules rules. The chart time range filters on each rule's creation date Rules that have never fired
Ingestion metrics ingestion. 365 days Log volume, freshness, parsing errors

 

2. Dashboard 1: Detection volume by severity

Metrics

  • Detections per day and per week, broken out by severity (for example CRITICAL, HIGH, MEDIUM, LOW)
  • Day-over-day and week-over-week change for each severity
  • Each severity's share of total volume, to show drift in the severity mix

Query: detections by severity over time

$date = timestamp.get_date(detection.created_time.seconds)$severity = detection.detection.severitymatch:  $date, $severityoutcome:  $detection_count = count_distinct(detection.id)order:  $date asc

Display it as a stacked bar or line chart with $date on the x-axis and the series split by $severity.

Gotcha: Don't add a detection.detection.severity != "" filter. It doesn't compile.

What to look for

Pattern Possible causes
CRITICAL spikes An active intrusion, a new or edited rule with a logic error, or a parser change creating extra UDM events
All severities decline together Telemetry loss, a rule disabled by accident, or a broken feed. Treat this as an outage until you've proven it isn't.
Severity mix shifts without a change in total volume Rules have been reclassified, or a curated rule set was enabled or changed

 

3. Dashboard 2: Rule performance

Metrics

  • Fire rate: detections per rule
  • Investigation rate and precision: taken from SOAR case closure data. This requires analysts to close cases with a consistent reason.
  • Rules that have never fired: enabled rules with no detections at all

Query A: top 10 rules by detection count

$rule_name = detection.detection.rule_namematch:  $rule_nameoutcome:  $count = count_distinct(detection.id)order:  $count desclimit:  10

Query B: enabled rules that have never fired

rules.live_status = "ENABLED"$rule_name = rules.name$display_name = rules.display_name$detection_time = rules.latest_detection_time.seconds$detection_time = 0match:  $rule_name, $display_name

Set this chart's time range to the maximum. The Rules data source applies the time range to each rule's creation date, so any rule created before the start of the range is hidden, even if it's enabled.

How to classify rules

Fire rate Precision Action
High Low Tune: add exclusions, tighten match windows, raise thresholds, or turn alerting off
High High Keep. Look into automating the response.
Low High Your best detections. Use them as templates for similar TTPs.
Never fired (Query B) n/a Check that the data it depends on exists. It may be a dead rule, or simply coverage for something rare.

Prerequisite: You can only measure precision if analysts record standardized dispositions, such as True Positive / Benign True Positive / False Positive. Without them, this dashboard can only show volume, not value.

4. Dashboard 3: Data source health

Metrics for each log type

  • Log volume per day, compared to a 7-day and 30-day baseline
  • Last log received, to spot sources that have gone quiet
  • Parsing errors, compared to log volume
  • Ingestion latency: Use the curated SecOps Log Monitoring dashboard, which already includes ingestion latency charts.

Query A: daily log count by log type

ingestion.component = "Ingestion API"$Log_Type = ingestion.log_type$Date = timestamp.get_date(ingestion.start_time)match:  $Date, $Log_Typeoutcome:  $Count = sum(ingestion.log_count)order:  $Date desc

Every log flow passes through the Ingestion API component, so filtering on it counts each log once.

Query B: last log received per log type

ingestion.component = "Ingestion API"$Log_Type = ingestion.log_typematch:  $Log_Typeoutcome:  $Recent_Ingestion_Time = timestamp.get_timestamp(max(ingestion.end_time), "%F %T")  $Total_Log_Volume = math.round(sum(ingestion.log_volume) / (1000 * 1000 * 1000), 2)order:  $Recent_Ingestion_Time asc

The stalest sources sort to the top, and $Total_Log_Volume is in GB. A log type that sent nothing during the whole time range won't appear at all, so compare the results against your list of expected sources.

Query C: parsing errors by log type

ingestion.component = "Normalizer"ingestion.state = "failed_parsing"$Log_Type = ingestion.log_typematch:  $Log_Typeoutcome:  $Parsing_Errors = sum(ingestion.log_count)order:  $Parsing_Errors desc
  • No rows: If Query A returns data but Query C returns no rows, no parsing failures were recorded in that time range.
  • Validation failures: Change the state to failed_validation and sum ingestion.event_count instead. Validation runs on parsed events, not raw logs.
  • Failure rate: Divide Query C's errors by Query A's log count for the same log type and time range.

Accuracy and access notes

  • Approximate counts: Ingestion metrics are aggregates sampled at intervals. Treat them as indicators, and use UDM search when you need exact event counts.
  • Data RBAC: Ingestion metrics follow your data access scope. If your scope includes a custom label (one built from a UDM regex or a data table), ingestion metrics are switched off for you entirely and these charts show no data. Scopes for ingestion monitoring should use only standard labels: log type, namespace, or ingestion source.
  • Ingestion source filters: If the dashboard is filtered by ingestion source, only log count populates. Byte and error charts can be empty, so filter by namespace instead.

Why it matters: Detection coverage depends on having the data. If EDR telemetry drops from 500k to 50k logs a day (−90%), your endpoint rules are still enabled but are effectively blind. No rule error will tell you this.

Pair the dashboard with alerting: Dashboards show trends, but someone has to look at them. To get notified, create Cloud Monitoring alerting policies on SecOps ingestion metrics. For example, a metric absence condition can fire when a log type or forwarder stops sending for 60 minutes. Health Hub also shows the status of each data source and parser.

5. Working out the cause of a trend

A change in volume is a symptom, not a verdict. Go through the possible causes before drawing a conclusion.

Volume rising:

  1. New or changed rule: Find the rule driving the increase with Dashboard 2 Query A. Then check when it was created or last updated (rules.create_time, rules.update_time).
  2. Parser change producing extra events: Check whether the parser or parser extension for that log type changed when the trend started. In ingestion metrics, if the ratio of Normalizer event_count to Ingestion API log_count rises for a log type, each raw log is producing more UDM events.
  3. Source expansion: New hosts, more users, or more verbose logging. Check Dashboard 3 Query A.
  4. Real increase in malicious activity: Only conclude this after ruling out 1–3.

Volume falling:

  1. Tuning worked: Check it against recent rule updates and exclusions.
  2. Source went silent or degraded: Check Dashboard 3 Queries A and B.
  3. Rule disabled or alerting turned off: Check rules.live_status and rules.alerting in the Rules data source.

6. Operational practice

  • Shift-start review: 5–10 minutes covering all three dashboards. Compare against the previous day (D-1) and the same day last week (D-7), so weekly cycles like weekends or batch jobs don't look like anomalies.
  • Baselines: For each metric, work out its normal daily range from the last 30 days, either the average plus or minus two standard deviations, or the middle 90% of values. Investigate any day that falls outside that range.
  • Ownership: Every anomaly gets an owner and a ticket. A dashboard nobody acts on is just decoration.
  • Weekly rule review: Use Dashboard 2 to pick the 10 noisiest rules (Query A) and all rules that have never fired (Query B) for tuning or retirement.
  • Playbook health: The Playbooks data source (playbook. prefix) records run status, so you can track playbook success and failure rates in the same review.
  • Start from curated content: Google publishes the queries behind its curated Data Ingestion and Health and SecOps Log Monitoring dashboards. Use them as a baseline for custom charts.

7. Action items

  • Build the three dashboards using the queries above
  • Set the time range on the never-fired rules chart to the maximum
  • Standardize case closure reasons so precision can be measured
  • Create Cloud Monitoring metric-absence alerts for critical log types
  • Add the dashboard review to the shift-start checklist
  • Set up a weekly review of the noisiest and never-fired rules