MON-004 Centralised Log Management
Description
Log data from production systems, applications, cloud services and network devices is aggregated into a centralised log management or SIEM platform. The platform provides search, correlation and reporting capability. Log ingestion coverage and health are monitored, and logging pipeline failures and ingestion gaps raise an alert to a responding team. The log management platform and the monitoring systems are themselves redundant, and their availability is monitored from outside the platform they run in, so the failure of the platform is reported by something the failure does not take down with it.
Rationale
Siloed logs across dozens of services are operationally unmanageable. Centralisation enables correlation of events across systems and is a prerequisite for effective threat detection. An ingestion-health dashboard inside the platform fails the second half by construction: the outage it is meant to catch is the one that stops it reporting. The independent watcher can be small, a heartbeat written and read from elsewhere, but it cannot run on the platform it is watching.
Applicability (9 profiles)
Annex point 3.2.6 is now stated: the log management platform and the monitoring systems are redundant and their availability is monitored from outside the platform they run in, so the outage that takes the platform down does not take the check with it.
Framework Mappings (20)
| LOG-01 | Logging and Monitoring Policy and Procedures | partial |
| LOG-03 | Security Monitoring and Alerting | informative |
| LOG-14 | Failures and Anomalies Reporting | partial |
| LOG-01 | Logging and Monitoring Policy and Procedures | partial |
| LOG-03 | Security Monitoring and Alerting | informative |
| LOG-14 | Failures and Anomalies Reporting | partial |
| HIPAA-164.308.a.1.ii.D | Information System Activity Review | informative |
| HIPAA-164.312.b | Audit Controls | informative |
| NIS2-CIR-3.2 | Monitoring and Logging | partial |
| NIS2-CIR-3.4 | Event Assessment and Classification | informative |
| AU-5 | Response to Audit Logging Process Failures | full |
| AU-6 | Audit Record Review, Analysis, and Reporting | informative |
| AU-6(1) | Audit Record Review, Analysis, and Reporting | Automated Process Integration | full |
| AU-6(3) | Audit Record Review, Analysis, and Reporting | Correlate Audit Record Repositories | full |
| AU-7 | Audit Record Reduction and Report Generation | full |
| AU-7(1) | Audit Record Reduction and Report Generation | Automatic Processing | full |
| CA-7 | Continuous Monitoring | informative |
| SI-4(1) | System Monitoring | System-wide Intrusion Detection System | full |
| SI-4(16) | System Monitoring | Correlate Monitoring Information | full |
| CC7.1 | Detection and Monitoring Procedures | partial |
Evidence (3)
SIEM or centralised log management platform configuration showing log source ingestion coverage across production systems, cloud services, applications, and network devices.
Example: Splunk, Elastic SIEM, AWS Security Lake, or equivalent platform configuration showing connected data sources, ingestion status per source, and last event received timestamp for each source
Test: Review the SIEM or log management platform data source inventory. Verify: (1) all production systems, cloud services, applications, and network devices appear as configured log sources; (2) each source shows a recent last-event-received timestamp (within expected interval); (3) ingestion health monitoring is enabled; (4) cross-reference the source list against the asset inventory to identify any ungapped systems.
Log ingestion health monitoring output showing pipeline status, ingestion volumes, and any detected gaps or failures in log collection.
Example: SIEM ingestion health dashboard export or monitoring alert configuration showing log source status, ingestion rate per source, and any sources with missed data in the last 30 days
Test: Query the log ingestion health dashboard for the last 30 days. Verify: (1) ingestion volume metrics are collected per log source; (2) pipeline failures or sources with zero ingestion trigger an alert; (3) any detected gaps have a documented investigation record; (4) coverage percentage for in-scope sources meets the defined threshold. (5) the platform's components are redundant, read from the deployment configuration rather than from a design document; (6) the availability of the platform is monitored by a system outside it, evidenced by an alert raised while the platform was unavailable, or by a test producing that result with its date.
Log storage capacity monitoring configuration showing threshold-based alerts for storage utilisation and logging pipeline health.
Example: CloudWatch alarm or Datadog monitor configuration showing storage capacity alert thresholds for log buckets and logging pipeline error rate alerts, with notification routing visible
Test: Review storage capacity monitoring configuration for all log storage locations. Verify: (1) capacity utilisation alerts are configured at a threshold that allows time for remediation before exhaustion; (2) logging pipeline failure or ingestion rate drop alerts are configured; (3) alerts route to an active response channel; (4) review the last 90 days of alerts to confirm alerts fired before any storage-related log loss.
Questions (3)
Is log data from production systems, applications, cloud services and network devices aggregated into a centralised log management platform?
Centralised collection is a prerequisite for effective threat detection. Siloed logs that cannot be correlated across systems leave blind spots in incident investigation.
Which of the following log sources are ingested into your centralised platform?
Options run from the most commonly ingested to the least. The control is about coverage and correlation, not about which platform is in use. A source that is collected but lands somewhere the platform cannot search does not count.
How are logging pipeline failures and log storage capacity issues detected and responded to?
Options run from strongest to weakest. Automated alerting with a defined response SLA and runbook is the working standard. The strongest option adds the part that is almost always missing: a check on the platform that runs somewhere else, because a health dashboard inside the platform goes dark with the outage it exists to report.