ThePlus Tech
Security operations · Case 02 of 03

Detection is easy. Closing the loop is the work.

Most tools produce signals. CyberGuardPlus turns signals into investigations that have an owner, a stage and an end, and keeps a hash-chained record of every privileged action taken along the way.

System
CyberGuardPlus
Role
Design, architecture and implementation
Shape
Kafka-staged modular monolith, Next.js console, Python sensor
Standing
Deployed and running on a UK host

The problem.

A small security team does not lack alerts. It lacks a defensible path from an alert to a closed investigation, and a record afterwards that survives someone asking what was done and by whom.

The wedge is lightweight explainable detection plus incident workflow for small teams. Signal quality first. Not an AI security operations centre replacement, and the repository says so in its own words to stop the scope drifting there.

The path through the system.

One sequence, in order. Each step exists because the next one cannot be trusted without it.

  1. 01Sensor agentA Python agent posts network and authentication events to the ingestion API.
  2. 02Raw topicsEvents land on Kafka before anything interprets them.
  3. 03NormalisationA consumer turns raw events into one normalised security event shape.
  4. 04Detection engineEvery detection rule runs against the normalised stream and emits detections.
  5. 05Alert scoringDetections are scored, then correlated into incidents rather than listed flat.
  6. 06InvestigationAn incident carries severity, stage, owner and lifecycle. A false positive is recorded as one, not deleted.
  7. 07Audit chainEvery privileged action is written to a hash-chained log that can be verified on demand.

The interface.

CyberGuardPlus incidents workspace listing seven investigations with severity, status, attack stages, alert counts and age.
Demo workspace, seeded data, signed in as demo.owner@cyberguard.local. One investigation is marked FALSE_POSITIVE and kept rather than removed.

What it is built from.

Core
Java · Spring Boot modular monolith
Pipeline
Kafka staged topics · Python sensor agent
Data
PostgreSQL · Redis · OpenSearch · MinIO
Console
Next.js 16 · React 19 · live-wired, no mock layer
Operations
Prometheus with 13 alert rules · Grafana · nginx with TLS
Deployment
Docker Compose on Ubuntu 24.04, OVHcloud Erith, United Kingdom

What was measured.

Every figure below is copied from this product’s own tracker, not written for this page.

CyberGuardPlus · evidenceRev 2026-09-08
Live deploymentTen core and monitoring containers running, aggregate health UP
HTTP 200
Demo-data isolationClean operator tenant returns zero incidents; sample tenant returns 4 incidents, 6 alerts, 3 sensors
0 vs 4
Restore drill, firstRestored into a throwaway database: 3 tenants, 2 users, 4 sensors, 4 incidents, 6 alerts, 45 audit rows
236,944 bytes
Restore drill, encrypted off-hostStreamed off the host, encrypted, restored: 8 incidents, 10 alerts, 46 audit rows, foreign keys and per-tenant audit continuity validated
Passed
Backup schedule30-day encrypted retention, run by launchd
Daily 03:15

What is not proven.

These stay listed until a measurement replaces them. They are part of the evidence, not a caveat attached to it.

  • This is a controlled pilot recovery posture, not production disaster recovery. The objective is RPO 24 hours and a target RTO of 4 hours, and the full destroy-and-rebuild duration has never been measured.
  • It is a production-grade reference scaffold, not a certified commercial security information and event management or endpoint detection product.
  • The outside-in health workflow on GitHub was refused before a runner started, because the account reports a billing block. A zero-cost probe from a local machine replaced it, which is a weaker instrument and is named as one.
  • No customer is using it. Both tenants above are operator-created.

What I would do next.

  1. 01Measure the destroy-and-rebuild time, which is the one number standing between controlled-pilot recovery and a real recovery-time objective.
  2. 02Restore a real continuous-integration probe rather than the local stand-in.
  3. 03Put it in front of one managed service provider and find out which part of the loop they would actually pay to close.

The other systems.

TrustLedger

What actually happened to the money?

Payment operations