AI Acceleration Impact
Project Duration (with AI)
Without AI
48d
With AI
34d
Saved
14d
Engineering Effort
Without AI
347h
With AI
233h
Saved
114h
Shared Plan · Read-only view. Sign in to move this plan to your workspace.
The AI Agent Observability & Trust Platform is being built for CISOs, AI platform leaders, and engineering teams at Fortune-500 companies. It addresses the need for real-time monitoring and management of enterprise AI agents by detecting behavioral drift and policy violations, while creating a structured failure memory to avoid repetitive error resolution. Key technical complexities include establishing a high-volume ingestion pipeline, integrating with various AI frameworks, and ensuring multi-tenancy and compliance across different deployment models.
Project Duration (with AI)
Without AI
48d
With AI
34d
Saved
14d
Engineering Effort
Without AI
347h
With AI
233h
Saved
114h
14
days (longest dependency chain)
Project duration: 34d · 3 sprints estimated
109
tasks
51 stories · 4 phases
Stories broken into tasks for execution
Foundation Setup
Establishes the foundational infrastructure, development environment, and initial project scaffolding.
Core Development
Focuses on implementing the core functionalities such as high-volume ingestion, real-time detection, and trace capture.
Integration and Governance
Integrates with external APIs and platforms, and develops governance features.
Advanced Features and Optimization
Implements advanced features like failure clustering, evaluation layers, and optimizes for scalability and compliance.
Establishes the foundational infrastructure, development environment, and initial project scaffolding.
Focuses on implementing the core functionalities such as high-volume ingestion, real-time detection, and trace capture.
Integrates with external APIs and platforms, and develops governance features.
Implements advanced features like failure clustering, evaluation layers, and optimizes for scalability and compliance.
This platform addresses the acute challenge of monitoring, governing, and ensuring the reliability and compliance of enterprise AI agents (e.g., LLM-based copilots, chatbots, and autonomous agents). As enterprises increasingly deploy AI agents in production, they face significant risks: agents can drift from intended behavior, violate policies, or repeatedly fail in unpredictable ways. These failures can lead to compliance breaches, security incidents, customer harm, and costly manual investigations. The problem is highly significant for Fortune-500 companies, where the stakes of AI misbehavior are regulatory, reputational, and financial. Current workarounds are fragmented: teams rely on ad hoc logging, manual incident review, basic output sampling, and custom scripts for drift detection—none of which scale or provide institutional memory of failures. There is no unified, real-time, enterprise-grade observability and trust layer for AI agents that meets the needs of CISOs, AI leaders, and engineering teams.
The solution provides a SaaS (and on-premise) API platform that ingests high-volume telemetry from AI agents, detects behavioral drift and policy violations in near real-time, and builds a 'failure memory'—a structured, searchable corpus of recurring agent errors. Key innovations include: (1) semantic drift detection using embeddings, (2) real-time and batch policy violation checks, (3) full LLM reasoning trace capture for forensic analysis, (4) automated clustering of non-deterministic failures using vector search and pattern discovery, and (5) role-based governance dashboards for both security and engineering. Compared to existing solutions (which are typically limited to basic logging, output sampling, or generic APM tools), this approach is purpose-built for the unique challenges of AI agent observability, compliance, and continuous evaluation at enterprise scale.
The platform offers enterprises the ability to deploy AI agents with confidence, knowing that behavioral drift, policy violations, and recurring failures will be detected and remediated before causing harm. Key benefits: (1) drastically reduced manual investigation time, (2) proactive risk and compliance management, (3) institutional memory of AI failures—so the same error is never solved twice, (4) faster and safer iteration on agent models, and (5) clear auditability for internal and external stakeholders. The unique selling proposition is the combination of real-time, semantic, and policy-aware observability with automated failure clustering and governance tailored for enterprise AI agents.
Primary users are CISOs, AI platform leaders, and engineering teams at large enterprises (especially Fortune-500 companies) deploying AI agents in production. These organizations operate in regulated industries (finance, healthcare, insurance, etc.) or have high reputational risk. Users are technical (developers, ML engineers) and security-focused (compliance, risk, and audit teams). The market size is substantial: as of 2026, over 60% of Fortune-500s have deployed LLM-based agents in customer-facing or internal workflows, with the enterprise AI observability and governance market estimated at $2.5B+ and growing at >30% CAGR. Key segments: regulated industries, tech-forward enterprises, and organizations with large-scale AI deployments.
The primary revenue model is SaaS subscription, priced by volume of agent telemetry/events ingested, number of monitored agents, and enterprise features (governance, compliance, on-premise deployment). For Fortune-500 customers, pricing typically ranges from $100K–$500K+ annually per enterprise, with pilots and design partnerships often starting at $25K–$50K for limited deployments. Additional revenue streams include professional services (integration, compliance audits), premium support, and usage-based overages for high-volume customers. Cost structure is dominated by cloud infrastructure (compute, storage for telemetry and vector DBs), LLM API usage, and engineering/DevOps. On-premise deployments may use a higher annual license plus support/maintenance fees.
high
mature
Research confidence: high
Research confidence: high
Threat Level: high
Confident AI provides an AI agent observability tool offering full trace visibility, step-by-step evaluations, research-backed metrics, human feedback integration, anomaly detection, and trace-to-dataset loops. It is designed for enterprise-grade AI agent deployments.
Strengths:
Weaknesses:
Sources:
Langfuse, acquired by ClickHouse in 2025, is an open-source platform for LLM observability, metrics, evaluations, prompt management, and experimentation, with a strong focus on data infrastructure.
Strengths:
Weaknesses:
Sources:
Helicone delivers proxy-based observability for AI agents, emphasizing minimal instrumentation overhead and rapid deployment for production environments.
Strengths:
Weaknesses:
Sources:
Braintrust is an 'evaluation-first' observability platform, focusing on structured evaluations and quality assurance within the AI agent observability loop.
Strengths:
Weaknesses:
Sources:
Maxim AI offers an end-to-end platform for AI simulation, evaluation, and observability, covering the full AI quality lifecycle from pre-release experimentation to real-time production monitoring.
Strengths:
Weaknesses:
Sources:
Chronosphere, now part of Palo Alto Networks, provides scalable observability for high data-volume and AI-driven environments, with integration into security and remediation workflows.
Strengths:
Weaknesses:
Sources:
Agent 365 is a Microsoft platform for deploying, securing, and monitoring AI agents, with features like telemetry, dashboards, and alerts, integrated with the Microsoft ecosystem.
Strengths:
Weaknesses:
Sources:
AgentSight is an observability framework introduced in 2025 that uses eBPF to monitor AI agents, aiming to bridge the semantic gap between high-level intent and low-level actions.
Strengths:
Weaknesses:
Sources:
Position as the only enterprise-grade platform that unifies high-volume agent telemetry ingestion, real-time semantic drift and policy violation detection, structured failure memory, and role-based governance dashboards with rapid deployment (within a day) and strict data isolation. Emphasize integration with leading LLM APIs (OpenAI, Anthropic) and compatibility with major agent frameworks (LangChain, custom orchestrators).
core-1Requirements:
Technical requirements:
core-2Requirements:
Technical requirements:
core-3Requirements:
Technical requirements:
core-4Requirements:
Technical requirements:
core-5Requirements:
Technical requirements:
core-6Requirements:
core-7Requirements:
Technical requirements:
core-8Requirements:
Technical requirements:
integration-1Requirements:
Integrations:
integration-2Requirements:
Integrations:
integration-3Requirements:
Integrations:
infra-1Requirements:
infra-2Requirements:
infra-3Requirements:
nfr-1Requirements:
nfr-2Requirements:
nfr-3Requirements:
nfr-4Requirements:
nfr-5Requirements:
Date: Saturday, August 1, 2026 Research as of: Saturday, August 1, 2026 (2026-08-01) Version: 1.0.0 Technology research confidence: high
Technology research sources:
The AI Agent Trust Monitor is a SaaS API service (with on-premise/air-gapped support) for ingesting, processing, and analyzing telemetry from enterprise AI agents. It provides real-time detection of behavioral drift and policy violations, maintains a structured "failure memory" for incident clustering, and supports interactive evaluation and governance workflows. The system is built on an event-driven microservices architecture with strict multi-tenancy, supporting high-throughput ingestion and low-latency alerting.
Key Components:
Data Flows:
System Boundaries:
Key Architectural Decisions:
flowchart TD
A["Agent-Side SDK"] --> B["Ingestion Gateway (FastAPI)"]
B --> C["Message Broker (Kafka-style)"]
C --> D1["Detection Service"]
C --> D2["Trace Repository"]
C --> D3["Evaluation Service"]
C --> D4["Failure Clustering Service"]
D1 --> E1["Alerting Service"]
D1 --> E2["Integration Connectors"]
D2 --> F1["Trace Search API"]
D4 --> F2["Failure Memory API"]
E1 --> G["Incident/Security Platforms"]
E2 --> G
H["Governance API (FastAPI)"] --> F1
H --> F2
subgraph "Observability & Monitoring"
Z1["OpenTelemetry Collector"]
Z2["Prometheus"]
Z3["Grafana"]
Z4["Centralized Logging"]
end
B --> Z1
D1 --> Z1
D2 --> Z1
D3 --> Z1
D4 --> Z1
E1 --> Z1
H --> Z1Primary Recommendations (one per category; version pins from web research):
| Category | Technology & Version | Rationale |
|---|---|---|
| Language/Runtime | Python 3.14.6 (docs) | Async-first, team expertise, latest stable, supports free-threaded/concurrency. |
| API Framework | FastAPI 0.140.0 (docs) | Async, OpenAPI-native, high performance, latest stable. |
| Message Broker | Apache Kafka 4.1 (Python aiokafka client) | Durable, high-throughput, partitioned, supports replay, proven at scale. |
| Task Queue | Celery 6 (with Redis 7 broker) | Async background/batch processing, Python-native, reliable. |
| Database (Transactional) | PostgreSQL 16 (with pgvector) | ACID, mature, multi-tenant, vector search for embeddings. |
| Search & Analytics | OpenSearch 3.0 | Fast, multi-tenant search for traces/failure memory, open-source, scalable. |
| In-Memory Cache | Redis 7 | Fast caching, pub/sub, queueing for low-latency. |
| Vector Search | pgvector (PostgreSQL extension) | Embedding-based similarity search for drift/failure clustering. |
| LLM API Integration | OpenAI API (2026), Anthropic API (2026) | Required for drift & eval. |
| Containerization | Docker 26 | Standard for deployability, reproducibility. |
| Orchestration | Docker Compose (Phase 1), Kubernetes 1.32 (Phase 2+) | Compose for solo/dev, K8s for scale/ops. |
| Observability | OpenTelemetry Collector 1.25, Prometheus 2.54, Grafana 11 |
Technology Compatibility:
Event-Driven Microservices:
All core features are implemented as separately deployable, async-first Python services, communicating via Kafka topics and REST/gRPC APIs.
Enables clear separation between fast-path (real-time detection) and heavy-path (batch eval, clustering).
CQRS (Command Query Responsibility Segregation):
Ingestion services write to event streams; query APIs read from optimized stores (OpenSearch, PostgreSQL).
Ensures ingestion is never blocked by slow queries.
Multi-Tenancy with Strict Isolation:
All data, processing, and API access are tenant-scoped by design, enforced at every layer (DB, broker, API).
Pattern Rationale:
Tables:
tenantstenant_configstelemetry_eventsevent_replay_queueSchema:
tenants
id (UUID, PK, DEFAULT gen_random_uuid())
name (VARCHAR(100), UNIQUE, NOT NULL)
status (ENUM: active, suspended, deleted, NOT NULL)
created_at (TIMESTAMP, DEFAULT NOW())
(INT, NOT NULL, DEFAULT 90)
Tables:
detectionsdetection_rulesdetection_alertsdetection_resultsSchema:
detections
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
telemetry_event_id (UUID, FK to telemetry_events.id, NOT NULL)
detection_type (ENUM: semantic_drift, policy_violation, anomaly, NOT NULL)
(FLOAT, NOT NULL)
Tables:
agent_tracestrace_stepstrace_metadataSchema:
agent_traces
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
agent_id (UUID, NOT NULL)
execution_id (VARCHAR(64), NOT NULL)
start_time (TIMESTAMP, NOT NULL)
Tables:
evaluationsevaluation_runsevaluation_resultsSchema:
evaluations
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
name (VARCHAR(100), NOT NULL)
mode (ENUM: offline, online, NOT NULL)
criteria (JSONB, NOT NULL)
Tables:
failure_clusterscluster_memberscluster_feedbackcluster_annotationsSchema:
failure_clusters
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
cluster_vector (VECTOR(1536), NOT NULL)
canonical_label (VARCHAR(128), NOT NULL)
(ENUM: active, merged, split, deleted, NOT NULL)
Tables:
usersrolesuser_rolesaudit_logsSchema:
users
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
email (VARCHAR(128), NOT NULL, UNIQUE)
password_hash (VARCHAR(128), NOT NULL)
(ENUM: active, suspended, deleted, NOT NULL)
Tables: (see tenant_configs, above)
Tables:
alert_integrationsalert_delivery_logsSchema:
alert_integrations
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
integration_type (ENUM: webhook, siem, api, pagerduty, splunk, servicenow, NOT NULL)
endpoint_url (VARCHAR(255), NOT NULL)
(JSONB, NULL)
erDiagram
tenants ||--o{ tenant_configs : "has"
tenants ||--o{ users : "has"
tenants ||--o{ agent_traces : "has"
tenants ||--o{ telemetry_events : "has"
tenants ||--o{ detections : "has"
tenants ||--o{ evaluations : "has"
tenants ||--o{ failure_clusters : "has"
tenants ||--o{ roles : "has"
tenants ||--o{ alert_integrations : "has"
users ||--o{ user_roles : "has"
roles ||--o{ user_roles : "has"
users ||--o{ cluster_feedback : "gives"
users ||--o{ cluster_annotations : "writes"
agent_traces ||--o{ trace_steps : "contains"
agent_traces ||--o{ trace_metadata : "has"
agent_traces ||--o{ cluster_members : "member"
agent_traces ||--o{ evaluation_results : "has"
telemetry_events ||--o{ detections : "generates"
telemetry_events ||--o{ event_replay_queue : "for"
detections ||--o{ detection_alerts : "triggers"
detections ||--o{ detection_results : "has"
detection_alerts ||--o{ alert_delivery_logs : "delivers"
failure_clusters ||--o{ cluster_members : "groups"
failure_clusters ||--o{ cluster_feedback : "receives"
failure_clusters ||--o{ cluster_annotations : "annotated"
evaluations ||--o{ evaluation_runs : "runs"
evaluation_runs ||--o{ evaluation_results : "produces"
alert_integrations ||--o{ alert_delivery_logs : "logs"
users ||--o{ audit_logs : "creates"Partitioning & Sharding:
telemetry_events, agent_traces, and failure_clusters partitioned by tenant_id and time for scale and retention.pgvector for fast similarity search./api/v1/POST /api/v1/telemetry/events
Request: { agent_id: str, event_type: str, event_timestamp: str, payload: dict }
Response: { id: str, status: "accepted" }
Auth: Bearer, tenant-scoped
Rate Limit: 10,000 req/min/tenant
GET /api/v1/telemetry/events?agent_id=:agent_id&from=:ts&to=:ts
Response:
GET /api/v1/detections?status=:status&type=:type
Response: [{ id, detection_type, score, threshold, detected_at, status }]
Auth: Bearer
POST /api/v1/detection/rules
Request: { rule_name: str, rule_type: str, criteria: dict, threshold: float }
Response: { id: str }
POST /api/v1/traces
Request: { agent_id: str, execution_id: str, steps: [ { step_type: str, step_payload: dict } ], metadata: dict }
Response: { id: str, status: "accepted" }
Auth: Bearer
GET /api/v1/traces/:trace_id
Response: { id, agent_id, execution_id, status, steps: [...], metadata: {...} }
POST /api/v1/evaluations
Request: { name: str, mode: "offline"|"online", criteria: dict, llm_judge_config?: dict, policy_check_config?: dict }
Response: { id: str }
Auth: Bearer, admin
POST /api/v1/evaluations/:eval_id/run
Request: { model_version: str }
GET /api/v1/failure_clusters
Response: [ { id, canonical_label, status, member_count } ]
Auth: Bearer
GET /api/v1/failure_clusters/:cluster_id
Response: { id, canonical_label, status, members: [...], annotations: [...] }
Auth: Bearer
GET /api/v1/tenants/me
Response: { id, name, compliance_policy, retention_policy_days }
Auth: Bearer
GET /api/v1/users/me
Response: { id, email, roles }
Auth: Bearer
GET /api/v1/audit_logs?resource_type=:type
GET /api/v1/tenant_configs
Response: { event_types, sampling_rate, policy_rules, semantic_thresholds }
Auth: Bearer, admin
PATCH /api/v1/tenant_configs
Request: { event_types?: list, sampling_rate?: float, policy_rules?: dict, semantic_thresholds?: dict }
Response: { status: "updated" }
POST /api/v1/alert_integrations
Request: { integration_type: str, endpoint_url: str, auth_config?: dict }
Response: { id: str }
Auth: Bearer, admin
GET /api/v1/alert_integrations
Response: [ { id, integration_type, endpoint_url, enabled } ]
POST /api/v1/users
Request: { email: str, password: str, roles: [str] }
Response: { id: str }
Auth: Bearer, admin
GET /api/v1/users
Response: [ { id, email, roles, status } ]
Auth: Bearer, admin
flowchart TD
A["Agent SDK"] --> B["/api/v1/telemetry/events (POST)"]
B --> C["Ingestion Gateway"]
C --> D["/api/v1/telemetry/events/replay (POST)"]
C --> E["/api/v1/detections (GET)"]
C --> F["/api/v1/traces (POST)"]
F --> G["/api/v1/traces/:trace_id (GET)"]
C --> H["/api/v1/evaluations (POST)"]
H --> I["/api/v1/evaluations/:eval_id/run (POST)"]
H --> J["/api/v1/evaluations/:eval_id/results (GET)"]
C --> K["/api/v1/failure_clusters (GET)"]
K --> L["/api/v1/failure_clusters/:cluster_id (GET)"]
L --> M["/api/v1/failure_clusters/:cluster_id/merge (POST)"]
L --> N["/api/v1/failure_clusters/:cluster_id/split (POST)"]
L --> O["/api/v1/failure_clusters/:cluster_id/annotate (POST)"]
L --> P["/api/v1/failure_clusters/:cluster_id/feedback (POST)"]
Q["/api/v1/alert_integrations (POST)"] --> R["Alerting Service"]
S["/api/v1/tenant_configs (PATCH)"] --> T["Tenant Config Service"]
U["/api/v1/auth/login (POST)"] --> V["Auth Service"]detections.detection_rules) per tenant/event.agent_traces/trace_steps, indexes vectors.failure_clusters.flowchart TD
SDK["AgentSDKAdapter"] --> IG["IngestionGatewayService"]
IG --> EB["EventBuffer (Kafka)"]
EB --> DP["DetectionProcessor"]
EB --> TC["TraceCaptureService"]
EB --> EO["EvaluationOrchestrator"]
DP --> DA["DetectionAlertDispatcher"]
DA --> AS["AlertingService"]
AS --> IC["IntegrationConnectors"]
TC --> TQ["TraceQueryService"]
EO --> EW["EvalWorker"]
EW --> LLM["LLMJudgeAdapter"]
EW --> FC["ClusteringService"]
FC --> CFM["ClusterFeedbackManager"]
IG --> TM["TenantIsolationMiddleware"]
IG --> AU["AuthService"]
AllServices --> TE["TelemetryExporter"]tenant_id; no cross-tenant access.Platform: Railway (managed PaaS, supports Docker Compose, managed PostgreSQL, Kafka, OpenSearch, Redis, Vault)
Phase 1 Topology:
Infrastructure-as-Code:
Dockerfile for each servicedocker-compose.yml with all services, networks, healthchecks.env file for secrets/config (never commit secrets; inject via Railway dashboard or Vault)Environment Variables:
DATABASE_URL (PostgreSQL)KAFKA_BROKER_URLOPENSEARCH_URLREDIS_URLVAULT_ADDRJWT_SECRETTENANT_ID (for on-prem/air-gapped)LLM_API_KEYS (for OpenAI/Anthropic, per-tenant or global)Database Setup:
pgvector)Deferred to Phase 2+:
Estimated Setup Time: 4–8 hours for a solo developer
Deployment Topology Diagram:
flowchart TD
subgraph "Railway Managed"
A1["IngestionGateway (FastAPI)"]
A2["DetectionService"]
A3["TraceService"]
A4["EvalService"]
A5["ClusteringService"]
A6["AlertingService"]
A7["GovernanceAPI"]
A8["AuthService"]
B1["PostgreSQL 16 (pgvector)"]
B2["Kafka 4.1"]
B3["OpenSearch 3.0"]
B4["Redis 7"]
B5["Vault 1.15"]
A1-->|"DB"|B1
A2-->|"DB"|B1
A3-->|"DB"|B1
A4-->|"DB"|B1
A5-->|"DB"|B1
A6-->|"DB"|B1
A7-->|"DB"|B1
A1-->|"Kafka"|B2
A2-->|"Kafka"|B2
A3-->|"Kafka"|B2
A4-->|"Kafka"|B2
A5-->|"Kafka"|B2
A6-->|"Kafka"|B2
A7-->|"Kafka"|B2
A1-->|"Search"|B3
A3-->|"Search"|B3
A5-->|"Search"|B3
A1-->|"Cache"|B4
A2-->|"Cache"|B4
A3-->|"Cache"|B4
A4-->|"Cache"|B4
A5-->|"Cache"|B4
A6-->|"Cache"|B4
A7-->|"Cache"|B4
A1-->|"Secrets"|B5
A2-->|"Secrets"|B5
A3-->|"Secrets"|B5
A4-->|"Secrets"|B5
A5-->|"Secrets"|B5
A6-->|"Secrets"|B5
A7-->|"Secrets"|B5
endMonolithic FastAPI App:
Simpler ops, but cannot scale ingestion and detection independently; not suitable for high-throughput, multi-tenant enterprise workloads.
RabbitMQ/Redis for Messaging:
Easier to manage, but lacks partitioning/replay required for enterprise event pipelines.
Single DB (PostgreSQL only):
Simpler, but search/analytics queries would degrade ingestion performance and limit fast vector search/clustering.
Serverless (Lambda/Step Functions):
Fast to prototype, but cold start/timeout limits and state management challenges for high-volume, low-latency ingestion.
When Alternatives Might Be Reconsidered:
Validation Checklist:
The AI Agent Observability & Trust Platform is an enterprise-grade SaaS (and on-premise/air-gapped) API service designed to monitor, govern, and ensure the reliability and compliance of AI agents (LLM-based copilots, chatbots, and autonomous agents) in Fortune-500 environments. The platform ingests high-volume telemetry from diverse agent frameworks, detects behavioral drift and policy violations in near real-time, and builds a structured, searchable "failure memory" to prevent repeated incident resolution. Key differentiators include semantic drift detection, automated failure clustering, full LLM reasoning trace capture, and role-based governance dashboards tailored for CISOs, AI platform leaders, and engineering teams. The solution is purpose-built for regulated, high-risk industries, providing rapid deployment, strict multi-tenancy, and seamless integration with leading LLM APIs and enterprise security tools.
Key Objectives:
As Fortune-500 enterprises increasingly deploy AI agents in production, they face acute challenges in monitoring, governing, and ensuring the reliability and compliance of these systems. Risks include behavioral drift, policy violations (e.g., data privacy, content compliance), and recurring, unpredictable failures that can lead to compliance breaches, security incidents, customer harm, and costly manual investigations. Existing workarounds—ad hoc logging, manual incident review, basic output sampling, and custom scripts—are fragmented, unscalable, and lack institutional memory. There is no unified, real-time, enterprise-grade observability and trust layer for AI agents that meets the needs of CISOs, AI leaders, and engineering teams, especially in regulated industries.
The platform provides a unified API service for ingesting, processing, and analyzing telemetry from enterprise AI agents. It features:
This approach delivers real-time, semantic, and policy-aware observability, actionable institutional memory, and rapid deployment (within a day), tailored for the unique needs of enterprise AI.
| Stakeholder | Role/Responsibility | Influence/Interest |
|---|---|---|
| CISOs/Security Leaders | Define compliance, monitor risk, audit incidents | High influence, high interest |
| AI Platform Leaders | Oversee agent deployment, ensure reliability | High influence, high interest |
| Engineering Teams (ML/DevOps) | Integrate SDK, respond to incidents, refine agents | High influence, high interest |
| Compliance & Audit Teams | Validate policy adherence, review audit logs | Medium influence, high interest |
| IT Operations | Manage deployment (SaaS/on-prem), monitor health | Medium influence, medium interest |
| Product Owner | Define requirements, prioritize roadmap | High influence, high interest |
| Executive Sponsors | Approve budget, set strategic direction | High influence, medium interest |
| Integration Partners | Connect incident/security tools, LLM APIs | Medium influence, medium interest |
All technical requirements are strictly aligned with the approved Architecture v1.0.0 (see reference above).
| Layer | Technology & Version |
|---|---|
| Language | Python 3.14.6 |
| API Framework | FastAPI 0.140.0 |
| Frontend | Angular (for dashboard, not included in API service scope) |
| Message Broker | Apache Kafka 4.1 (aiokafka client) |
| Task Queue | Celery 6 (with Redis 7) |
| Database | PostgreSQL 16 (with pgvector extension) |
| Search | OpenSearch 3.0 |
| In-Memory | Redis 7 |
| Vector Search | pgvector (PostgreSQL extension) |
| LLM APIs | OpenAI API (2026), Anthropic API (2026) |
| Container | Docker 26 |
| Orchestration | Docker Compose (Phase 1), Kubernetes 1.32 (Phase 2+) |
| Observability | OpenTelemetry Collector 1.25, Prometheus 2.54, Grafana 11 |
| Secrets Mgmt | HashiCorp Vault 1.15 |
| AuthN/AuthZ | JWT (FastAPI), OAuth2, RBAC middleware |
| API Docs | OpenAPI 3.1 (FastAPI auto-gen), Swagger UI |
| Object Store | S3-compatible (MinIO for on-prem, AWS S3 for SaaS) |
| Category | Requirement |
|---|---|
| Performance | Sub-second to low-second latency for real-time detection; batch checks within minutes-hours. |
| Scalability | Support for 10,000+ concurrent agents per tenant; horizontally scalable ingestion/detection. |
| Reliability | 99.9% uptime (SaaS); auto-retry on failed LLM API calls; DLQ for failed events. |
| Availability | Multi-AZ (SaaS); on-premise supports HA via Docker Compose/K8s. |
| Security | TLS 1.3, AES-256-at-rest, RBAC, audit logs, DDoS protection, Vault for secrets. |
| Privacy | Strict per-tenant isolation; GDPR/CCPA support; data export/delete APIs. |
| Compliance | Out-of-the-box support for GDPR, CCPA, HIPAA, FINRA; audit trails; configurable retention. |
| Observability | OpenTelemetry traces, Prometheus metrics, Grafana dashboards, centralized logging. |
| Maintainability | OpenAPI contracts, modular microservices, CI/CD pipelines, codegen SDKs. |
| Deployability | SaaS (managed), on-premise/air-gapped (Docker Compose/K8s), <1 day onboarding. |
| Metric | Target/Goal | Measurement Method |
|---|---|---|
| Time to Detect Critical Drift/Policy Violation | <2 seconds (real-time tier) | Synthetic test agents, alert logs |
| Reduction in Manual Incident Investigation Time | >50% reduction vs. baseline | User surveys, incident tracking |
| Agent Onboarding Time (SDK integration) | <1 day from install to first event | Onboarding analytics |
| Repeat Incident Rate (Same Failure Twice) | <5% of incidents are repeats | Failure memory cluster analytics |
| Compliance Auditability (Log Completeness) | 100% of actions/audits logged per tenant | Audit log sampling |
| Uptime (SaaS) | 99.9% | Prometheus/Grafana monitoring |
| Tenant Data Isolation Breaches | 0 | Security incident logs |
| Integration Success Rate (Alert Delivery) | >99% successful delivery to external tools | Alert delivery logs |
| Risk Category | Description | Mitigation/Contingency |
|---|---|---|
| Technical | Scaling ingestion and detection at enterprise telemetry volumes | Kafka partitioning, horizontal scaling, back-pressure, DLQ |
| Market | Enterprises may prefer in-house or incumbent APM solutions | Emphasize unique features, rapid deployment, integration flexibility |
| Competitive | Large APM/logging vendors or cloud providers may enter the space | Focus on AI agent-specific features, failure memory, rapid onboarding |
| Execution | Complexity of robust multi-tenancy, compliance, and on-premise support from day one | Strict architecture enforcement, automated tests, phased rollout |
| LLM API Dependency | Outages or API changes in OpenAI/Anthropic | Graceful degradation, local fallback, on-prem proxy, batch reprocessing |
| Security | Data leakage or cross-tenant access | RBAC, tenant-scoped queries, regular pen-testing, audit logging |
| Integration | Failure to deliver alerts to customer tools | Retry logic, DLQ, alert delivery logs, manual resend tools |
tenants, tenant_configs, telemetry_events, event_replay_queue, detections, detection_rules, detection_alerts, detection_results, agent_traces, trace_steps, , , , , , , , , , , , , , .telemetry_events| Column | Type | Description |
|---|---|---|
| id | UUID, PK | Event ID |
| tenant_id | UUID, FK | Tenant isolation |
| agent_id | UUID | Agent identifier |
| event_type | ENUM | tool_call, reasoning_step, state_transition, response, error |
| event_timestamp | TIMESTAMP | Event occurrence time |
| payload | JSONB | Structured event data |
| ingested_at | TIMESTAMP | Ingestion time |
| replay_status | ENUM | pending, processed, failed |
erDiagram
tenants ||--o{ tenant_configs : "has"
tenants ||--o{ users : "has"
tenants ||--o{ agent_traces : "has"
tenants ||--o{ telemetry_events : "has"
tenants ||--o{ detections : "has"
tenants ||--o{ evaluations : "has"
tenants ||--o{ failure_clusters : "has"
tenants ||--o{ roles : "has"
tenants ||--o{ alert_integrations : "has"
users ||--o{ user_roles : "has"
roles ||--o{ user_roles : "has"
users ||--o{ cluster_feedback : "gives"
users ||--o{ cluster_annotations : "writes"
agent_traces ||--o{ trace_steps : "contains"
agent_traces ||--o{ trace_metadata : "has"
agent_traces ||--o{ cluster_members : "member"
agent_traces ||--o{ evaluation_results : "has"
telemetry_events ||--o{ detections : "generates"
telemetry_events ||--o{ event_replay_queue : "for"
detections ||--o{ detection_alerts : "triggers"
detections ||--o{ detection_results : "has"
detection_alerts ||--o{ alert_delivery_logs : "delivers"
failure_clusters ||--o{ cluster_members : "groups"
failure_clusters ||--o{ cluster_feedback : "receives"
failure_clusters ||--o{ cluster_annotations : "annotated"
evaluations ||--o{ evaluation_runs : "runs"
evaluation_runs ||--o{ evaluation_results : "produces"
alert_integrations ||--o{ alert_delivery_logs : "logs"
users ||--o{ audit_logs : "creates"All endpoints are under /api/v1/ and strictly match the Architecture.
POST /api/v1/telemetry/events
{ agent_id: str, event_type: str, event_timestamp: str, payload: dict }{ id: str, status: "accepted" }GET /api/v1/telemetry/events?agent_id=:agent_id&from=:ts&to=:ts
[{ id, agent_id, event_type, event_timestamp, payload }]POST /api/v1/telemetry/events/replay
{ event_ids: [str] }GET /api/v1/detections?status=:status&type=:type
[{ id, detection_type, score, threshold, detected_at, status }]POST /api/v1/detection/rules
{ rule_name: str, rule_type: str, criteria: dict, threshold: float }{ id: str }GET /api/v1/detection/rules
[{ id, rule_name, rule_type, criteria, threshold, enabled }]POST /api/v1/traces
{ agent_id: str, execution_id: str, steps: [ { step_type: str, step_payload: dict } ], metadata: dict }{ id: str, status: "accepted" }GET /api/v1/traces/:trace_id
{ id, agent_id, execution_id, status, steps: [...], metadata: {...} }GET /api/v1/traces/search?query=:q
[{ id, agent_id, execution_id, status, start_time }]POST /api/v1/evaluations
{ name: str, mode: "offline"|"online", criteria: dict, llm_judge_config?: dict, policy_check_config?: dict }{ id: str }POST /api/v1/evaluations/:eval_id/run
{ model_version: str }{ run_id: str, status: "started" }GET /api/v1/evaluations/:eval_id/runs
GET /api/v1/failure_clusters
[ { id, canonical_label, status, member_count } ]GET /api/v1/failure_clusters/:cluster_id
{ id, canonical_label, status, members: [...], annotations: [...] }POST /api/v1/failure_clusters/:cluster_id/merge
{ target_cluster_id: str }{ status: "merged" }GET /api/v1/tenants/me
{ id, name, compliance_policy, retention_policy_days }GET /api/v1/users/me
{ id, email, roles }GET /api/v1/audit_logs?resource_type=:type
[ { id, action, resource_type, resource_id, timestamp } ]POST /api/v1/alert_integrations
{ integration_type: str, endpoint_url: str, auth_config?: dict }{ id: str }GET /api/v1/alert_integrations
[ { id, integration_type, endpoint_url, enabled } ]PATCH /api/v1/alert_integrations/:integration_id
{ enabled?: bool, endpoint_url?: str, auth_config?: dict }POST /api/v1/users
{ email: str, password: str, roles: [str] }{ id: str }GET /api/v1/users
[ { id, email, roles, status } ]PATCH /api/v1/users/:user_id
{ status?: str, roles?: [str] }{ status: "updated" }All endpoints require JWT/OAuth2 Bearer tokens, are tenant-scoped, and enforce RBAC.
/health endpoints.| Phase | Deliverables | Timeline | Dependencies |
|---|---|---|---|
| Phase 1 | Core ingestion, detection, trace, clustering, API | 8 weeks | Architecture, SDK |
| Phase 2 | Governance dashboard, advanced integrations, K8s | 6 weeks | Phase 1 |
| Phase 3 | Multi-region SaaS, blue-green, advanced analytics | 8 weeks | Phase 2 |
End of PRD
This epic covers providing secure authentication and authorization mechanisms for all API endpoints, SDKs, and integrations to protect tenant data and platform integrity, including role-based access control and secure secrets management.
As the system, I want to enforce RBAC at API and data layers, so that users and services have access only to authorized resources within their tenant scope.
Priority: high
Dependencies: Implement JWT and OAuth2 Authentication for APIs and SDKs (37ae5097-ac24-4534-8d8b-c323d39eafd1)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As the system, I want to authenticate all API and SDK requests using JWT Bearer tokens and OAuth2 flows, so that only authorized users and services can access platform resources.
Priority: high
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers implementing configurable data retention policies and compliance controls per tenant to meet enterprise regulatory requirements such as GDPR, CCPA, HIPAA, and FINRA.
As the system, I want to automatically delete or export tenant data according to retention policies and compliance requests, so that regulatory requirements like GDPR and CCPA are met.
Priority: high
Dependencies: Implement Tenant APIs for Data Retention Configuration (3a796e08-9a77-4884-8941-b5832a0982a9)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As a tenant admin, I want APIs to specify data retention durations and deletion policies, so that platform usage complies with legal and regulatory obligations.
Priority: high
Dependencies: Implement Configurable Data Retention and Compliance Policies (1ea10a17-3f8b-4bf8-ac20-ac6d66d63f08)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
This epic covers ensuring the platform can scale horizontally to handle high-volume telemetry ingestion, processing, storage, and querying for large enterprise tenants while maintaining performance under peak loads.
As the platform, I want to optimize database and search engine performance for large telemetry and trace data volumes, so that queries remain performant at scale.
Priority: high
Dependencies: Implement Trace Indexing and Search API (38b5afff-c3aa-4a82-bdee-b815a12a1422), Implement Embedding-Based Failure Clustering Service (240dba66-daf2-40f4-8744-620f0a605707)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As the platform, I want ingestion and processing services to scale horizontally, so that the system can handle 10,000+ concurrent agents per tenant without degradation.
Priority: high
Dependencies: Implement Kafka-Based Event Buffer with Tenant Partitioning (665e5481-0184-48ad-a0b6-1feb018c5607), Implement Async Microservices for Detection, Trace, Evaluation, and Clustering (74545937-1014-43f9-a5ff-7ff049d89d7f)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers supporting highly flexible policy violation detection rules that allow customers to define custom rules, adjust thresholds, and extend semantic/contextual checks via APIs or user interfaces, enabling tailored compliance monitoring.
As a tenant admin, I want to create, modify, and extend policy violation detection rules via API, so that I can customize compliance monitoring to evolving regulations and internal policies.
Priority: high
Dependencies: Develop APIs for Managing Detection Rules (db361ded-1f1b-4c02-99e7-a681bef98703)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
This epic covers ensuring that critical behavioral drift and policy violation alerts are generated and delivered with sub-second to low-second latency to enable proactive incident response, while batch checks operate on longer intervals for trend analysis.
As the detection system, I want to run batch detection checks on longer intervals (minutes to hours) for trend analysis, so that comprehensive evaluation is possible without impacting real-time performance.
Priority: high
Dependencies: Implement Behavioral Anomaly Scoring with Tiered Latency Targets (03282f1e-057f-4b62-8d49-e0e2a0968dc6)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
As the detection and alerting system, I want to generate and deliver critical alerts within sub-second to low-second latency, so that customers can respond proactively to AI agent issues.
Priority: high
Dependencies: Implement Behavioral Anomaly Scoring with Tiered Latency Targets (03282f1e-057f-4b62-8d49-e0e2a0968dc6), Implement Alert Delivery via Webhooks, SIEM, and APIs (c982f1d3-9ef3-4ebf-9a1d-b402338e0fa2)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers the implementation of a microservices architecture with an event-driven design using a message broker to separate fast ingestion paths from heavier asynchronous processing. It ensures scalability and fault tolerance of ingestion and processing pipelines.
As the platform, I want to implement microservices for detection, trace capture, evaluation, and failure clustering that consume events asynchronously from Kafka, so that processing is scalable and decoupled.
Priority: high
Dependencies: Implement Kafka-Based Event Buffer with Tenant Partitioning (665e5481-0184-48ad-a0b6-1feb018c5607)
Acceptance Criteria:
Story Points: 13
Estimated Effort: 13 hours
As the ingestion gateway, I want to enqueue telemetry events to a Kafka message broker partitioned by tenant, so that ingestion is scalable, isolated, and supports replay.
Priority: high
Dependencies: Build High-Throughput Ingestion Gateway Service (29b7d171-303b-423f-8e61-2db646a7f49b)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers deployment flexibility to support fully managed SaaS with continuous updates and integrations, as well as on-premise or air-gapped deployments with limited or manual integration updates and stricter data flow controls.
As the platform maintainer, I want to support on-premise and air-gapped deployments with local proxies for external integrations and manual update mechanisms, so that customers with strict data flow controls can use the platform securely.
Priority: high
Dependencies: Design SaaS Deployment with Managed Integrations and Rapid Onboarding (db913a78-4926-4c88-b7d1-838957041e6c)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As the platform maintainer, I want to provide a SaaS deployment model with managed integrations and rapid onboarding, so that customers can quickly deploy and scale the platform.
Priority: high
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers the implementation of strict per-tenant data isolation and retention policies to meet enterprise compliance and security requirements. It ensures logical separation and access control of tenant data across all system layers.
As a tenant admin, I want to specify data retention durations and deletion policies, so that platform usage aligns with legal and regulatory obligations.
Priority: high
Dependencies: Implement Tenant Data Partitioning and Access Control (4dd60acc-5193-4d8b-80d4-b0774bf05d33)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As the system, I want to enforce tenant data partitioning and access control at database and API layers, so that tenant data privacy and compliance are guaranteed.
Priority: high
Dependencies: Build High-Throughput Ingestion Gateway Service (29b7d171-303b-423f-8e61-2db646a7f49b)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers integration with enterprise incident management and security tools such as PagerDuty, Splunk, and ServiceNow. It supports alert delivery using standard protocols and configurable routing per tenant.
As a tenant admin, I want APIs to configure which incident management tools to integrate with and how alerts are routed, so that alerting fits organizational processes.
Priority: high
Dependencies: Implement Integration Connectors for PagerDuty, Splunk, and ServiceNow (9bab17cb-f56e-4c6a-a966-d54cea7f89bd)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
As the integration connector, I want to deliver alerts to PagerDuty, Splunk, and ServiceNow using their APIs and protocols, so that alerts are integrated into customer workflows.
Priority: high
Dependencies: Implement Alert Delivery via Webhooks, SIEM, and APIs (c982f1d3-9ef3-4ebf-9a1d-b402338e0fa2)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers integration with the Anthropic API to support AI/LLM layer functionalities including semantic drift detection, evaluation, and reasoning trace capture. It includes authentication management and error handling.
As the LLMJudgeAdapter, I want to call Anthropic API for embeddings, evaluation, and trace enrichment, so that the platform provides complementary AI capabilities.
Priority: high
Dependencies: Implement Offline Evaluation Framework with Historical Failure Memory (5d292973-057d-4626-b9ce-cdcedd3ed7ae), Implement Embedding-Based Semantic Drift Detection (72840e4e-e601-42e8-b1eb-ad9e85d0a506)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
This epic covers integration with the OpenAI API to support AI/LLM layer functionalities such as semantic drift detection, evaluation scoring, and reasoning trace enrichment. It includes handling authentication, rate limits, and error handling.
As the LLMJudgeAdapter, I want to call OpenAI API to generate embeddings and perform evaluation scoring, so that the platform leverages advanced AI capabilities for semantic analysis.
Priority: high
Dependencies: Implement Offline Evaluation Framework with Historical Failure Memory (5d292973-057d-4626-b9ce-cdcedd3ed7ae), Implement Embedding-Based Semantic Drift Detection (72840e4e-e601-42e8-b1eb-ad9e85d0a506)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
This epic covers integration of alerting mechanisms with customers’ existing incident management and security platforms via standard protocols. It supports configurable alert routing, role-based notifications, and secure, reliable delivery to enable proactive incident response.
As a tenant admin, I want APIs to configure alert integrations, enable/disable endpoints, and set routing preferences per tenant and user role, so that alerting aligns with organizational workflows.
Priority: high
Dependencies: Implement Alert Delivery via Webhooks, SIEM, and APIs (c982f1d3-9ef3-4ebf-9a1d-b402338e0fa2)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
As the alerting service, I want to deliver alerts to external incident management and security platforms using webhooks, SIEM connectors, and APIs, so that alerts integrate seamlessly into enterprise workflows.
Priority: high
Dependencies: Build Alerting Mechanisms Integrated with Incident Management Tools (d6aa9452-5767-4d55-8541-cb25c40dcfe4)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers tenant control over telemetry data collection via the SDK, including configuration of event types, sampling rates, and policy violation detection rules or semantic thresholds. It ensures observability is tailored to compliance and operational needs without disrupting agent workflows.
As a tenant admin, I want to define and extend policy violation detection rules and semantic thresholds via SDK configuration, so that detection is tailored to evolving compliance requirements.
Priority: high
Dependencies: Expose Tenant Configuration APIs for Telemetry Event Types and Sampling (1d28e93c-1db7-4655-b6d0-7a70f6a0ad36)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
As a tenant admin, I want APIs to configure which telemetry event types to capture and adjust sampling rates, so that I can balance observability with performance and compliance.
Priority: high
Dependencies: Implement Tenant-Specific Telemetry Configuration APIs in SDK (67259a43-c948-4905-8525-e11bab7a4fe7)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
This epic covers the development of role-based governance dashboards tailored for CISOs/security leaders and engineering teams. It provides views for policy compliance, audit trails, risk perimeter, execution trace exploration, regression alignment, and remediation tracking.
As an engineering team member, I want dashboard views showing execution traces, regression test results, and remediation status, so that I can investigate and resolve AI agent issues efficiently.
Priority: medium
Dependencies: Implement Role-Based Access Control for Governance API (051f3a23-007e-4b08-bd97-2c5a79ce3b8b), Implement Trace Indexing and Search API (38b5afff-c3aa-4a82-bdee-b815a12a1422), Implement Embedding-Based Failure Clustering Service (240dba66-daf2-40f4-8744-620f0a605707)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As a CISO, I want dashboard views showing policy compliance status, audit trails, and risk perimeter metrics, so that I can monitor organizational AI agent governance effectively.
Priority: medium
Dependencies: Implement Role-Based Access Control for Governance API (051f3a23-007e-4b08-bd97-2c5a79ce3b8b)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As the system, I want to enforce role-based access control on governance API endpoints, so that users see data relevant to their roles and permissions.
Priority: medium
Dependencies: Build High-Throughput Ingestion Gateway Service (29b7d171-303b-423f-8e61-2db646a7f49b)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
This epic covers the automatic grouping of non-deterministic production incidents into structured canonical failure modes using embedding similarity and pattern discovery. It supports user-in-the-loop feedback, manual overrides, annotations, and maintains a searchable failure memory per tenant.
As an engineering team member, I want to validate, merge, split, and annotate failure clusters via APIs, so that failure taxonomies evolve with organizational knowledge.
Priority: high
Dependencies: Implement Embedding-Based Failure Clustering Service (240dba66-daf2-40f4-8744-620f0a605707)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As the clustering service, I want to automatically group recurring AI failures into canonical failure modes using embedding similarity and pattern discovery, so that engineering teams get a structured failure knowledge base.
Priority: high
Dependencies: Implement Trace Indexing and Search API (38b5afff-c3aa-4a82-bdee-b815a12a1422)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers the evaluation framework that supports offline regression testing of new AI model versions against historical failure memory and continuous online scoring of live agent outputs for policy compliance and quality. It integrates LLM-as-judge and rule-based policy checks with configurable criteria per tenant.
As a tenant admin, I want APIs to create, run, and query evaluation jobs and results, so that I can manage evaluation workflows and monitor outcomes.
Priority: high
Dependencies: Implement Offline Evaluation Framework with Historical Failure Memory (5d292973-057d-4626-b9ce-cdcedd3ed7ae)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
As the system, I want to continuously score live production agent outputs for policy compliance and quality baselines, so that ongoing compliance is monitored in real-time.
Priority: high
Dependencies: Implement Offline Evaluation Framework with Historical Failure Memory (5d292973-057d-4626-b9ce-cdcedd3ed7ae), Implement Embedding-Based Semantic Drift Detection (72840e4e-e601-42e8-b1eb-ad9e85d0a506)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As an engineering team member, I want to run offline regression tests of new AI model versions against historical failure memory, so that I can validate model quality before deployment.
Priority: high
Dependencies: Implement Embedding-Based Failure Clustering Service (240dba66-daf2-40f4-8744-620f0a605707), Implement Trace Capture API and Storage (8aa13535-ea85-4533-b219-3ef970a4a2da)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers capturing and storing full execution traces of AI agent runs, including tool call sequences, intermediate reasoning steps, and contextual metadata. It supports multi-tenant isolation, indexing, and efficient retrieval for forensic analysis and replay.
As the system, I want to ensure strict multi-tenant isolation of trace data at storage and API layers, so that tenant data privacy and compliance are maintained.
Priority: high
Dependencies: Implement Trace Capture API and Storage (8aa13535-ea85-4533-b219-3ef970a4a2da)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
As a maintainer, I want to index trace data with vector embeddings and provide search APIs, so that users can efficiently retrieve and replay execution traces.
Priority: high
Dependencies: Implement Trace Capture API and Storage (8aa13535-ea85-4533-b219-3ef970a4a2da)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As a library consumer, I want to submit detailed reasoning traces including steps and metadata via API, so that full execution context of AI agent runs is captured for analysis.
Priority: high
Dependencies: Build High-Throughput Ingestion Gateway Service (29b7d171-303b-423f-8e61-2db646a7f49b)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers the implementation of a detection system that performs semantic drift detection, policy violation checks, and behavioral anomaly scoring with tiered latency targets. It includes configurable detection rules, alerting mechanisms, and integration with incident management tools to enable proactive monitoring of AI agent behavior.
As a tenant integrator, I want APIs to create, read, update, and disable detection rules, so that I can tailor policy violation detection to evolving compliance needs.
Priority: high
Dependencies: Build High-Throughput Ingestion Gateway Service (29b7d171-303b-423f-8e61-2db646a7f49b)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
As the alerting service, I want to deliver detection alerts via webhooks, SIEM connectors, and APIs integrated with customer incident management platforms, so that alerts reach the right teams promptly.
Priority: high
Dependencies: Implement Behavioral Anomaly Scoring with Tiered Latency Targets (03282f1e-057f-4b62-8d49-e0e2a0968dc6)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As the detection processor, I want to score behavioral anomalies with tiered latency targets, providing sub-second alerts for critical issues and batch processing for less urgent checks, so that alerting is timely and scalable.
Priority: high
Dependencies: Implement Embedding-Based Semantic Drift Detection (72840e4e-e601-42e8-b1eb-ad9e85d0a506), Develop Policy Violation Detection Engine with Configurable Rules (89b1a2ee-c1cc-4044-9ac2-d81bd5d526b2)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As the detection processor, I want to detect policy violations including data privacy, content compliance, and security policies using configurable rules per tenant, so that compliance issues are identified in real-time.
Priority: high
Dependencies: Build High-Throughput Ingestion Gateway Service (29b7d171-303b-423f-8e61-2db646a7f49b), Implement Per-Tenant Data Isolation and Retention Policies in Ingestion Pipeline (86893a21-1b1e-4ac6-826d-c7be40dc3d8a)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As the detection processor, I want to perform embedding-based semantic drift detection comparing live telemetry against baseline agent behavior per tenant, so that behavioral deviations are detected early.
Priority: high
Dependencies: Build High-Throughput Ingestion Gateway Service (29b7d171-303b-423f-8e61-2db646a7f49b), Implement Per-Tenant Data Isolation and Retention Policies in Ingestion Pipeline (86893a21-1b1e-4ac6-826d-c7be40dc3d8a)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
This epic covers the design and implementation of a scalable ingestion pipeline that allows agent-side SDKs to push structured telemetry events at enterprise scale. It includes SDK modularity, tenant-specific configuration, data isolation, and replay capabilities to support reliable telemetry ingestion without disrupting existing AI agent workflows.
As a maintainer, I want to support replaying historical telemetry events on demand, so that downstream processing can reprocess events for updated detection or evaluation logic.
Priority: high
Dependencies: Build High-Throughput Ingestion Gateway Service (29b7d171-303b-423f-8e61-2db646a7f49b), Implement Per-Tenant Data Isolation and Retention Policies in Ingestion Pipeline (86893a21-1b1e-4ac6-826d-c7be40dc3d8a)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
As the ingestion pipeline, I want to enforce strict per-tenant data isolation and configurable retention policies, so that tenant data privacy and compliance requirements are met.
Priority: high
Dependencies: Build High-Throughput Ingestion Gateway Service (29b7d171-303b-423f-8e61-2db646a7f49b)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As the ingestion gateway service, I want to accept telemetry events from SDKs asynchronously, authenticate tenants, validate payloads, and enqueue events to a Kafka message broker with per-tenant partitioning, so that high-volume telemetry is ingested reliably and isolated per tenant.
Priority: high
Dependencies: Develop Modular Agent-Side SDK with Framework Adapters (79865b34-da6c-4fbd-8b91-d0c743499707)
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
As a tenant integrator, I want to configure telemetry data collection via SDK APIs to select event types, adjust sampling rates, and define policy rules, so that telemetry observability aligns with compliance and operational needs.
Priority: high
Dependencies: Develop Modular Agent-Side SDK with Framework Adapters (79865b34-da6c-4fbd-8b91-d0c743499707)
Acceptance Criteria:
Story Points: 5
Estimated Effort: 5 hours
As a library consumer (developer), I want a modular agent-side SDK with adapters for diverse AI agent frameworks like LangChain and custom orchestrators, so that I can integrate telemetry event pushing seamlessly without disrupting existing workflows.
Priority: high
Acceptance Criteria:
Story Points: 8
Estimated Effort: 8 hours
Spike
As a technical writer, I want to set up documentation frameworks and generate initial API docs using OpenAPI and Swagger UI, so that developers have access to up-to-date API specifications.
Priority: medium
Timebox: 3 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 3
Spike
As a developer, I want to set up code quality tools like ESLint and Prettier for frontend and static analysis tools for backend, so that code consistency and quality are maintained.
Priority: medium
Timebox: 2 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 2
Spike
As a QA engineer, I need to set up testing frameworks (pytest for backend, Jest/Vitest for frontend), so that automated tests can be written and executed efficiently.
Priority: high
Timebox: 3 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 3
Spike
As a backend engineer, I need to set up a database migration framework (e.g., Alembic) for PostgreSQL, so that schema changes can be versioned and applied consistently.
Priority: high
Timebox: 3 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 3
Spike
As a DevOps engineer, I need to configure CI/CD pipelines using GitHub Actions for build, test, and deployment automation, so that code changes are validated and deployed reliably.
Priority: high
Timebox: 5 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 5
Spike
As a developer, I need to set up a local development environment using Docker Compose and necessary dev tools, so that I can develop and test services locally.
Priority: high
Timebox: 3 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 3
Spike
As a developer, I need to initialize the project repository and scaffold the initial project structure, so that development can start with a clean, organized codebase.
Priority: high
Timebox: 3 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 3
Execution order should follow dependencies; complete "Depends on" tasks before each task.
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
0f868508-4e96-42bd-908b-dc0f1d071dd1Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
0f868508-4e96-42bd-908b-dc0f1d071dd1Acceptance Criteria:
Definition of Done:
2172ecd6-17c5-42e9-8dac-1f555195f672Acceptance Criteria:
Definition of Done:
fb464dab-cf46-472a-b80d-558a660c2f00, cc26a7ab-bee1-4763-aca4-9be547ab8a6c, a1faf49e-fb14-4769-8b1d-089f7416d91fAcceptance Criteria:
Definition of Done:
a1faf49e-fb14-4769-8b1d-089f7416d91fAcceptance Criteria:
Definition of Done:
e3a6d381-c7d6-4967-a5f3-36395356fa59Acceptance Criteria:
Definition of Done:
a1faf49e-fb14-4769-8b1d-089f7416d91f, e7f64d84-eb40-4fb1-a598-63d65f943d4bAcceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
2b1471fc-0a1f-4229-980d-48e99adb9958Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
77f334c9-1412-4ec7-8d6c-4ddd8ced05f2Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
e330ee3a-d325-445f-966a-96e25f9973cbAcceptance Criteria:
Definition of Done:
3f2ee986-5fc8-4774-96bb-ce21c2e103ae, 8e4a4607-56d7-4715-a9e0-acbd6f1dbe7bAcceptance Criteria:
Definition of Done:
77923311-08f5-4e0c-b045-d738d97447f0Acceptance Criteria:
Definition of Done:
108d341c-102c-4e2b-8d33-95ff23f62182, bdc41ad7-d29d-4164-92e8-2892a454e781, a1faf49e-fb14-4769-8b1d-089f7416d91fAcceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
8bde14fd-a1fd-4866-969e-ad57498becf4Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
ec88b461-6f5c-44e4-873d-36f682f674b6Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
e330ee3a-d325-445f-966a-96e25f9973cb, ec88b461-6f5c-44e4-873d-36f682f674b6, 2502336e-d013-4850-838b-96d1d5fe3501, 39e52612-2c8a-4668-88e0-ef5b2b8a885dAcceptance Criteria:
Definition of Done:
7a3bb596-06af-4400-ba64-de216aaa69e4Acceptance Criteria:
Definition of Done:
ec88b461-6f5c-44e4-873d-36f682f674b6Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
b245d527-aede-449a-86d7-1c35e8000acdAcceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
ed86bbf4-cf7d-4a82-8936-2a18f980690d, dde6b355-b610-43de-be63-74e0a48f5e32Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
ed86bbf4-cf7d-4a82-8936-2a18f980690dAcceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
3b39227b-921e-4c1e-b6cc-b8e90959e528Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
90e91dcc-6857-4449-b707-88e1f84be11bAcceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
d798526c-b2a8-47b0-b674-9c620ed62968Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
9ae3e6ea-8be0-47a5-8ab3-fbe8caa49f00Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
932afadc-dba0-4fc0-8690-9dad690c9b47Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
dcdad24d-d870-41bb-b4fd-c7c848c6d0c3Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
e7f64d84-eb40-4fb1-a598-63d65f943d4bAcceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
ec88b461-6f5c-44e4-873d-36f682f674b6, 9d29eb82-98ae-48f8-ac35-c5058537e715Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
7c9bb509-d909-4e2f-8384-9274d2666f07Acceptance Criteria:
Definition of Done:
3b39227b-921e-4c1e-b6cc-b8e90959e528Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
| Metrics, logs, traces, dashboards. |
| Secrets Management | HashiCorp Vault 1.15 | Centralized, secure secret storage. |
| CI/CD | GitHub Actions | Integrated with source control, flexible pipelines. |
| AuthN/AuthZ | JWT (FastAPI), OAuth2, RBAC middleware | Secure, multi-tenant, enterprise-compliant. |
| API Docs/Contracts | OpenAPI 3.1 (FastAPI auto-gen), Swagger UI | Standardized, developer-friendly. |
| File/Object Storage | S3-compatible (MinIO for on-prem, AWS S3 for SaaS) | Scalable, flexible for on-prem/SaaS. |
OpenAPI-First:
All APIs are defined with OpenAPI contracts, enabling SDK/codegen and external auditability.
Observability-First:
OpenTelemetry traces, metrics, and logs are emitted from all services, supporting enterprise monitoring.
retention_policy_dayscompliance_policy (JSONB, NOT NULL)
INDEX idx_tenants_name ON tenants(name)
tenant_configs
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
event_types (JSONB, NOT NULL) // e.g., ["tool_call", "reasoning_step"]
sampling_rate (FLOAT, NOT NULL, DEFAULT 1.0)
policy_rules (JSONB, NOT NULL)
semantic_thresholds (JSONB, NOT NULL)
updated_at (TIMESTAMP, DEFAULT NOW())
UNIQUE (tenant_id)
INDEX idx_tenant_configs_tenant_id ON tenant_configs(tenant_id)
telemetry_events
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
agent_id (UUID, NOT NULL)
event_type (ENUM: tool_call, reasoning_step, state_transition, response, error, NOT NULL)
event_timestamp (TIMESTAMP, NOT NULL)
payload (JSONB, NOT NULL)
ingested_at (TIMESTAMP, DEFAULT NOW())
replay_status (ENUM: pending, processed, failed, DEFAULT NULL)
INDEX idx_events_tenant_id_type ON telemetry_events(tenant_id, event_type)
INDEX idx_events_agent_id ON telemetry_events(agent_id)
PARTITION BY RANGE (event_timestamp) (for scale/retention)
event_replay_queue
id (UUID, PK, DEFAULT gen_random_uuid())
telemetry_event_id (UUID, FK to telemetry_events.id, NOT NULL)
replay_requested_at (TIMESTAMP, DEFAULT NOW())
replay_status (ENUM: pending, processing, completed, failed, NOT NULL)
INDEX idx_replay_status ON event_replay_queue(replay_status)
scorethreshold (FLOAT, NOT NULL)
detected_at (TIMESTAMP, DEFAULT NOW())
status (ENUM: new, acknowledged, resolved, ignored, NOT NULL)
INDEX idx_detections_tenant_type ON detections(tenant_id, detection_type)
INDEX idx_detections_status ON detections(status)
detection_rules
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
rule_name (VARCHAR(100), NOT NULL)
rule_type (ENUM: semantic, policy, anomaly, NOT NULL)
criteria (JSONB, NOT NULL)
threshold (FLOAT, NOT NULL)
enabled (BOOLEAN, DEFAULT TRUE)
created_at (TIMESTAMP, DEFAULT NOW())
INDEX idx_rules_tenant_id ON detection_rules(tenant_id)
detection_alerts
id (UUID, PK, DEFAULT gen_random_uuid())
detection_id (UUID, FK to detections.id, NOT NULL)
alert_type (ENUM: webhook, siem, api, NOT NULL)
destination (VARCHAR(255), NOT NULL)
alert_payload (JSONB, NOT NULL)
sent_at (TIMESTAMP, DEFAULT NOW())
status (ENUM: pending, sent, failed, NOT NULL)
INDEX idx_alerts_status ON detection_alerts(status)
detection_results
id (UUID, PK, DEFAULT gen_random_uuid())
detection_id (UUID, FK to detections.id, NOT NULL)
result_payload (JSONB, NOT NULL)
evaluated_at (TIMESTAMP, DEFAULT NOW())
end_time (TIMESTAMP)
status (ENUM: success, error, timeout, NOT NULL)
trace_vector (VECTOR(1536), NULL) // Embedding for similarity search
created_at (TIMESTAMP, DEFAULT NOW())
INDEX idx_traces_tenant_agent ON agent_traces(tenant_id, agent_id)
INDEX idx_traces_vector ON agent_traces USING ivfflat(trace_vector) // pgvector
trace_steps
id (UUID, PK, DEFAULT gen_random_uuid())
trace_id (UUID, FK to agent_traces.id, NOT NULL)
step_index (INT, NOT NULL)
step_type (ENUM: tool_call, reasoning, output, error, NOT NULL)
step_payload (JSONB, NOT NULL)
created_at (TIMESTAMP, DEFAULT NOW())
UNIQUE (trace_id, step_index)
trace_metadata
id (UUID, PK, DEFAULT gen_random_uuid())
trace_id (UUID, FK to agent_traces.id, NOT NULL)
key (VARCHAR(64), NOT NULL)
value (VARCHAR(256), NOT NULL)
UNIQUE (trace_id, key)
llm_judge_config (JSONB, NULL)
policy_check_config (JSONB, NULL)
created_at (TIMESTAMP, DEFAULT NOW())
INDEX idx_evaluations_tenant_mode ON evaluations(tenant_id, mode)
evaluation_runs
id (UUID, PK, DEFAULT gen_random_uuid())
evaluation_id (UUID, FK to evaluations.id, NOT NULL)
run_timestamp (TIMESTAMP, DEFAULT NOW())
model_version (VARCHAR(64), NOT NULL)
status (ENUM: running, completed, failed, NOT NULL)
INDEX idx_eval_runs_eval_id ON evaluation_runs(evaluation_id)
evaluation_results
id (UUID, PK, DEFAULT gen_random_uuid())
run_id (UUID, FK to evaluation_runs.id, NOT NULL)
trace_id (UUID, FK to agent_traces.id, NOT NULL)
score (FLOAT, NOT NULL)
result_payload (JSONB, NOT NULL)
evaluated_at (TIMESTAMP, DEFAULT NOW())
INDEX idx_eval_results_run_id ON evaluation_results(run_id)
INDEX idx_eval_results_trace_id ON evaluation_results(trace_id)
statuscreated_at (TIMESTAMP, DEFAULT NOW())
INDEX idx_clusters_tenant_label ON failure_clusters(tenant_id, canonical_label)
INDEX idx_clusters_vector ON failure_clusters USING ivfflat(cluster_vector)
cluster_members
id (UUID, PK, DEFAULT gen_random_uuid())
cluster_id (UUID, FK to failure_clusters.id, NOT NULL)
trace_id (UUID, FK to agent_traces.id, NOT NULL)
added_at (TIMESTAMP, DEFAULT NOW())
UNIQUE (cluster_id, trace_id)
cluster_feedback
id (UUID, PK, DEFAULT gen_random_uuid())
cluster_id (UUID, FK to failure_clusters.id, NOT NULL)
user_id (UUID, FK to users.id, NOT NULL)
feedback_type (ENUM: validate, merge, split, NOT NULL)
details (JSONB, NULL)
created_at (TIMESTAMP, DEFAULT NOW())
cluster_annotations
id (UUID, PK, DEFAULT gen_random_uuid())
cluster_id (UUID, FK to failure_clusters.id, NOT NULL)
user_id (UUID, FK to users.id, NOT NULL)
annotation_text (TEXT, NOT NULL)
created_at (TIMESTAMP, DEFAULT NOW())
statuscreated_at (TIMESTAMP, DEFAULT NOW())
INDEX idx_users_tenant_email ON users(tenant_id, email)
roles
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
role_name (VARCHAR(64), NOT NULL)
permissions (JSONB, NOT NULL)
UNIQUE (tenant_id, role_name)
user_roles
id (UUID, PK, DEFAULT gen_random_uuid())
user_id (UUID, FK to users.id, NOT NULL)
role_id (UUID, FK to roles.id, NOT NULL)
UNIQUE (user_id, role_id)
audit_logs
id (UUID, PK, DEFAULT gen_random_uuid())
tenant_id (UUID, FK to tenants.id, NOT NULL)
user_id (UUID, FK to users.id, NOT NULL)
action (VARCHAR(128), NOT NULL)
resource_type (VARCHAR(64), NOT NULL)
resource_id (UUID, NOT NULL)
timestamp (TIMESTAMP, DEFAULT NOW())
details (JSONB, NULL)
INDEX idx_audit_tenant_resource ON audit_logs(tenant_id, resource_type)
auth_configenabled (BOOLEAN, DEFAULT TRUE)
created_at (TIMESTAMP, DEFAULT NOW())
INDEX idx_alert_integrations_tenant_type ON alert_integrations(tenant_id, integration_type)
alert_delivery_logs
id (UUID, PK, DEFAULT gen_random_uuid())
alert_integration_id (UUID, FK to alert_integrations.id, NOT NULL)
alert_id (UUID, FK to detection_alerts.id, NOT NULL)
delivery_status (ENUM: pending, delivered, failed, NOT NULL)
delivery_attempted_at (TIMESTAMP, DEFAULT NOW())
response_code (INT, NULL)
response_body (TEXT, NULL)
[{ id, agent_id, event_type, event_timestamp, payload }]Auth: Bearer
POST /api/v1/telemetry/events/replay
Request: { event_ids: [str] }
Response: { replay_job_id: str }
Auth: Bearer, admin
Auth: Bearer, admin
GET /api/v1/detection/rules
Response: [{ id, rule_name, rule_type, criteria, threshold, enabled }]
Auth: Bearer
PATCH /api/v1/detection/rules/:rule_id
Request: { enabled: bool, threshold?: float, criteria?: dict }
Response: { id: str }
Auth: Bearer, admin
POST /api/v1/detections/:detection_id/acknowledge
Response: { id: str, status: "acknowledged" }
Auth: Bearer
Auth: Bearer
GET /api/v1/traces/search?query=:q
Response: [{ id, agent_id, execution_id, status, start_time }]
Auth: Bearer
Response: { run_id: str, status: "started" }
Auth: Bearer, admin
GET /api/v1/evaluations/:eval_id/runs
Response: [ { run_id, status, run_timestamp, model_version } ]
Auth: Bearer
GET /api/v1/evaluations/:eval_id/results
Response: [ { trace_id, score, result_payload, evaluated_at } ]
Auth: Bearer
POST /api/v1/failure_clusters/:cluster_id/mergeRequest: { target_cluster_id: str }
Response: { status: "merged" }
Auth: Bearer, admin
POST /api/v1/failure_clusters/:cluster_id/split
Request: { member_trace_ids: [str] }
Response: { new_cluster_id: str }
Auth: Bearer, admin
POST /api/v1/failure_clusters/:cluster_id/annotate
Request: { annotation_text: str }
Response: { id: str }
Auth: Bearer
POST /api/v1/failure_clusters/:cluster_id/feedback
Request: { feedback_type: "validate"|"merge"|"split", details?: dict }
Response: { id: str }
Auth: Bearer
Response: [ { id, action, resource_type, resource_id, timestamp } ]
Auth: Bearer, admin
Auth: Bearer, admin
Auth: Bearer
PATCH /api/v1/alert_integrations/:integration_id
Request: { enabled?: bool, endpoint_url?: str, auth_config?: dict }
Response: { status: "updated" }
Auth: Bearer, admin
GET /api/v1/alert_delivery_logs?status=:status
Response: [ { id, alert_id, delivery_status, delivery_attempted_at, response_code } ]
Auth: Bearer, admin
PATCH /api/v1/users/:user_id
Request: { status?: str, roles?: [str] }
Response: { status: "updated" }
Auth: Bearer, admin
POST /api/v1/auth/login
Request: { email: str, password: str }
Response: { access_token: str, refresh_token: str }
Auth: None
POST /api/v1/auth/refresh
Request: { refresh_token: str }
Response: { access_token: str }
Auth: None
No OpenTelemetry/Observability:
Would limit production readiness and troubleshooting in regulated environments.
trace_metadataevaluationsevaluation_runsevaluation_resultsfailure_clusterscluster_memberscluster_feedbackcluster_annotationsusersrolesuser_rolesaudit_logsalert_integrationsalert_delivery_logs{ replay_job_id: str }PATCH /api/v1/detection/rules/:rule_id
{ enabled: bool, threshold?: float, criteria?: dict }{ id: str }POST /api/v1/detections/:detection_id/acknowledge
{ id: str, status: "acknowledged" }[ { run_id, status, run_timestamp, model_version } ]GET /api/v1/evaluations/:eval_id/results
[ { trace_id, score, result_payload, evaluated_at } ]POST /api/v1/failure_clusters/:cluster_id/split
{ member_trace_ids: [str] }{ new_cluster_id: str }POST /api/v1/failure_clusters/:cluster_id/annotate
{ annotation_text: str }{ id: str }POST /api/v1/failure_clusters/:cluster_id/feedback
{ feedback_type: "validate"|"merge"|"split", details?: dict }{ id: str }GET /api/v1/tenant_configs{ event_types, sampling_rate, policy_rules, semantic_thresholds }PATCH /api/v1/tenant_configs
{ event_types?: list, sampling_rate?: float, policy_rules?: dict, semantic_thresholds?: dict }{ status: "updated" }{ status: "updated" }GET /api/v1/alert_delivery_logs?status=:status
[ { id, alert_id, delivery_status, delivery_attempted_at, response_code } ]POST /api/v1/auth/login
{ email: str, password: str }{ access_token: str, refresh_token: str }POST /api/v1/auth/refresh
{ refresh_token: str }{ access_token: str }