CloudTrail Threat Detection Platform
Real-time Security Event Ingestion & Threat Analysis
1. SYSTEM OVERVIEW
An event-driven cloud security analytics pipeline that ingests AWS CloudTrail records, applies normalization and risk scoring heuristics, and fires alert integrations under 45 seconds to secure distributed multi-tenant AWS accounts.
Security teams were blind to privilege escalations, impossible travel access anomalies, and critical token abuses, with incident containment cycles taking hours due to fragmented SIEM reporting.
Reduced MTTR (Mean Time to Respond) from 3 hours to under 3 minutes, mitigated risk of credentials leakage, and established audit-ready continuous cloud monitoring compliance.
- Ingest and normalize 10k event/sec spikes without queue blockages
- Filter out 95% of safe system noise via context-aware suppression
- Deploy immutable evidence stores for post-incident audit retention
- Designed least-privilege IAM control models for distributed log acquisition
- Implemented Lambda scoring heuristics engine and EventBridge routing
- Engineered sub-minute alert pathways into ChatOps and incident workflows
2. SYSTEM FEATURES
Serverless S3-triggered event pipelines normalizing raw JSON CloudTrail schemas.
Mathematical incident prioritization engine scoring API calls against historical baselines.
ChatOps alert integration formatting threat metadata with direct links to logs.
3. SYSTEM ARCHITECTURE
Distributed security architecture acquiring audit logs from organizational member accounts into a centralized security account using encrypted streams, then processing events asynchronously via concurrent serverless queues.
CloudTrail Logs (S3) ──> S3 Event Notification ──> Lambda Parser ──> EventBridge
│
┌──────────────────────────────────────────────────────────────────────┘
▼
Lambda Risk Scorer ──> DynamoDB (State Cache) ──> KMS Encrypted Evidence (S3)
│
└─[High Risk Risk >= 75]─> SNS ──> ChatOps (Slack) / PagerDuty Alertcloudtrail-detector/
├── src/
│ ├── handlers/
│ │ ├── parser.ts # Normalization logic
│ │ └── scorer.ts # Risk scoring engine
│ ├── rules/
│ │ └── privilege-escalation.ts # Signature checks
│ └── utils/
│ └── kms-helper.ts # Evidence encryption
└── terraform/
├── main.tf # Pipelines definitions
└── variables.tf4. INTERACTIVE SIMULATOR WIDGET
Run active operations audits utilizing the custom sandbox telemetry receiver widget below.
5. ENGINEERING ARCHITECTURE DECISIONS (ADRs)
Log streams arrive in bursts and processing must be real-time while ensuring zero dropouts.
- Managed Kafka (MSK)
- Kinesis Data Streams
- Serverless S3 Event Notifications
Serverless S3 Event Notifications triggering Lambda
- Zero idle cost during low activity periods
- Scales automatically to match arbitrary incoming log volumes
- Extremely simple architecture with minimal operational footprint
- Subject to Lambda cold starts
- Slightly higher latency compared to persistent consumers
Accepted a 1-second cold start latency on initial batch wakeups to achieve 85% cost savings compared to maintaining a running MSK cluster.
Introduce an SQS buffer layer between S3 and Lambda to smooth out heavy burst workloads and handle retry logic.
6. DETAILED TECHNOLOGY STACK
7. SECURITY REVIEW & POSTURE
Centralized security parsing functions must query DynamoDB and read logs across organizational accounts without global administrative permissions.
Intruders attempting to cover tracks might try to modify or delete logs and threat records in the evidence store.
8. PERFORMANCE METRICS TELEMETRY
Optim:Direct event routing, batch size optimizations, and memory tuning of parser Lambdas.
Optim:Configured concurrent execution limits and partitioned DynamoDB keys for risk caches.
9. ENGINEERING CHALLENGES & RESOLUTIONS
High volumes of benign system activity (e.g. deployment roles creating access keys) triggered false positives, drowning security teams in warnings.
Root Cause:Heuristics were too simple, matching only on high-severity API calls without examining the calling entity context or history.
Traced alert history over 30 days and observed that 88% of alerts originated from trusted automated CI/CD roles executing predictable tasks.
Implemented context-aware suppression. Roles that matched a specific cryptographically signed identity configuration and ran inside known IP CIDR blocks were automatically suppressed or assigned a lower risk score.
Introduced some latency in evaluation to check the DynamoDB identity directory, adding ~40ms to processing time.
Static rules always decay; threat detection systems must adapt to environment contexts and identify normal automated patterns.
10. SYSTEM LESSONS LEARNED
Clean log ingestion requires strong schema validation on entry, as API structures occasionally shift without notice.
Asynchronous event decoupling is essential; parsing failures must never block downstream incident processing.
Operational security metrics (false-positive ratios, analyst review times) are just as critical as architectural ones.
I would replace the DynamoDB caching layer with Redis to lower evaluation latency and enable easier cluster-wide scale.
11. ROADMAP & TECHNICAL DEBT
- Integrate LLM-assisted alert summaries for security analysts
- Implement active-response automated IAM isolation loops
- Support Google Cloud audit logs ingestion
12. GITHUB SOURCE EXPLORER
Audit raw repository script configurations directly inside the active terminal workspace.