AI Acceleration Impact
Project Duration (with AI)
Without AI
52d
With AI
39d
Saved
13d
Engineering Effort
Without AI
935h
With AI
686h
Saved
249h
Shared Plan · Read-only view. Sign in to move this plan to your workspace.
The project is a cloud operating system designed for organizations seeking an efficient and integrated solution for managing cloud infrastructure. It addresses the limitations of traditional virtualization stacks by enabling direct-to-metal operations, which reduces overhead and improves performance. The system integrates essential cloud services and allows for rapid deployment of Kubernetes clusters and containers. Key technical challenges include ensuring seamless synchronization of state across distributed nodes and developing a robust operations layer that supports natural language commands for infrastructure management.
Project Duration (with AI)
Without AI
52d
With AI
39d
Saved
13d
Engineering Effort
Without AI
935h
With AI
686h
Saved
249h
16
days (longest dependency chain)
Project duration: 39d · 3 sprints estimated
70
tasks
40 stories · 5 phases
Stories broken into tasks for execution
Foundation Setup
Establishes the foundational infrastructure, including CI/CD, development environment, and initial project scaffolding.
Core Development
Focuses on developing the core functionalities of the system, including the Kubernetes and Container Engines.
Integration and Testing
Integrates core components and conducts comprehensive testing to ensure system reliability and performance.
Advanced Features and Optimization
Implements advanced features such as AI-native operations and optimizes system performance.
Deployment and Launch
Finalizes deployment strategies and prepares the system for launch.
Establishes the foundational infrastructure, including CI/CD, development environment, and initial project scaffolding.
Focuses on developing the core functionalities of the system, including the Kubernetes and Container Engines.
Integrates core components and conducts comprehensive testing to ensure system reliability and performance.
Implements advanced features such as AI-native operations and optimizes system performance.
Finalizes deployment strategies and prepares the system for launch.
qbOS addresses the complexity, inefficiency, and fragmentation of current cloud infrastructure management. Traditional cloud environments rely on layered virtualization stacks (hypervisors, VMs, then Kubernetes or containers), which introduce significant overhead, operational complexity, and performance penalties. Operators must stitch together disparate infrastructure services (networking, firewall, DNS, storage, etc.), often from multiple vendors, leading to integration pain, increased attack surface, and slow provisioning. The problem is significant for organizations running private/hybrid clouds, edge deployments, or large-scale on-premises infrastructure—especially those seeking bare-metal performance, simplified management, and strong security. Current workarounds include using converged infrastructure platforms (e.g., VMware Cloud Foundation, OpenStack), DIY Kubernetes on bare metal (with complex manual integration), or relying on public cloud managed services (with less control and higher costs). None offer a truly integrated, direct-to-metal, AI-native experience.
qbOS eliminates the virtualization layer by running directly on hardware, integrating all critical infrastructure services into a single operating system. The ORIGIN and LINK model enables seamless, distributed cloud formation without external control planes or cluster federation, reducing complexity and improving resilience. The AI-native, prompt-driven operations layer (using MCP and AsyncAPI) allows operators to manage infrastructure via natural language, automating tasks that typically require deep expertise and manual scripting. Key innovations include: direct-to-metal container execution (QCE), rapid Kubernetes cluster provisioning (QKE), unified distributed state sync, and a fully integrated, AI-driven operations interface. Compared to existing solutions, qbOS offers lower overhead, faster provisioning, tighter integration, and a fundamentally new user experience.
qbOS delivers a unified, high-performance cloud operating system that radically simplifies infrastructure management by eliminating virtualization overhead, integrating all core services, and enabling AI-native operations. Key benefits: faster provisioning (bare-metal clusters in under 60 seconds), lower operational complexity, improved security (integrated firewall, access control, TLS), direct hardware access for demanding workloads, and a natural-language-driven interface that reduces skill barriers. The unique selling proposition is its direct-to-metal, fully integrated, AI-native architecture—unlike competitors that bolt AI onto legacy stacks or require complex multi-layer setups.
Primary users are cloud infrastructure operators, DevOps teams, and IT administrators in enterprises, SaaS providers, research institutions, and edge computing vendors. Demographics include mid-to-large enterprises with private/hybrid cloud needs, organizations with high-performance or GPU workloads (AI/ML, HPC), and service providers seeking to simplify infrastructure management. The market size is substantial: the global private/hybrid cloud infrastructure market exceeded $100B in 2025, with strong growth in edge and AI infrastructure. Key segments: enterprise IT, AI/ML infrastructure, edge/cloud service providers, and regulated industries (finance, healthcare, government) requiring on-premises control.
qbOS is delivered as self-hosted software, with revenue generated through annual enterprise licensing, per-node or per-core subscriptions, and premium support/consulting. Additional revenue streams may include managed updates, training, and integration services. Pricing strategies should reflect the value of operational simplification and performance gains, with tiered offerings for different organization sizes and workloads. (No explicit pricing data provided in source material; recommend validation with design partners.)
high
mature
Research confidence: high
Research confidence: medium
Threat Level: medium
Hadron Linux, launched in January 2026, is a lightweight, security-focused Linux distribution optimized for enterprise edge and cloud deployments. It is designed for immutable, minimal installations and is part of the CNCF's Kairos project.
Strengths:
Weaknesses:
Sources:
Equinix Metal provides bare-metal cloud services, granting users direct access to physical servers without virtualization. It has a global data center presence and a strong API-driven provisioning model.
Strengths:
Weaknesses:
Sources:
IBM Cloud offers a range of bare-metal instances, including specialized Z and LinuxONE mainframe architectures, targeting workloads that require unique hardware or compliance.
Strengths:
Weaknesses:
Sources:
Talos Linux is an immutable, minimal operating system purpose-built for Kubernetes clusters, emphasizing security and operational simplicity with an API-driven approach.
Strengths:
Weaknesses:
Sources:
Flatcar is a community-driven, immutable OS designed for container-centric and Kubernetes environments, offering automated updates and consistent behavior.
Strengths:
Weaknesses:
Sources:
Containerd is a lightweight, industry-standard container runtime, widely used as the core engine for container orchestration platforms like Kubernetes.
Strengths:
Weaknesses:
Sources:
Position qbOS as the only AI-native, self-contained cloud operating system that can instantly provision production-grade Kubernetes and containers directly on bare metal, with integrated infrastructure services and a unified distributed architecture that eliminates the need for traditional virtualization and orchestration stacks. Emphasize operational simplicity, security, and speed, targeting organizations seeking to modernize or simplify their cloud infrastructure.
core-1Requirements:
Technical requirements:
core-2Requirements:
Technical requirements:
core-3Requirements:
Technical requirements:
core-4Requirements:
Technical requirements:
core-5Requirements:
Technical requirements:
core-6Requirements:
integration-1Requirements:
Integrations:
infra-1Requirements:
nfr-1Requirements:
nfr-2Requirements:
nfr-3Requirements:
Date: Wednesday, July 29, 2026 Research as of: Wednesday, July 29, 2026 (2026-07-29) Version: 1.0.0 Technology research confidence: high
Technology research sources:
DirectMetal Cloud System (qbOS) is a self-hosted, direct-to-metal cloud operating system for operators seeking to deploy and control cloud infrastructure without traditional virtualization. It integrates all critical cloud services—networking, firewall, DNS, ACME/TLS, load balancing, storage, container registry, and access control—into a single, unified OS. The system is anchored by two core execution layers:
The architecture is fundamentally distributed, using an ORIGIN and LINK node model. An ORIGIN node initializes the cloud, and any number of LINK nodes can be added, with all nodes synchronizing state directly using a custom, low-latency distributed consensus protocol (Raft/Paxos inspired but tailored for qbOS).
A prompt-driven, AI-native operations layer—using MCP and AsyncAPI—enables natural-language control of all major infrastructure and orchestration functions.
Key Architectural Decisions:
flowchart TD
A["Operator Interface (Web/CLI, Prompt-driven)"]
B["AI-Native Operations Layer (MCP + AsyncAPI)"]
C["Distributed State Layer (Consensus, Metadata, Policy)"]
D1["ORIGIN Node"]
D2["LINK Node(s)"]
E1["Kubernetes Engine (QKE)"]
E2["Container Engine (QCE)"]
F1["Integrated Networking Service"]
F2["Integrated Firewall Service"]
F3["Integrated DNS Service"]
F4["ACME/TLS Service"]
F5["Load Balancer Service"]
F6["Storage Service"]
F7["Container Registry"]
F8["Access Control Service"]
G["External Integrations (ACME, Container Registry, Hardware Drivers)"]
A --> B
B --> C
C <--> D1
C <--> D2
D1 <--> D2
D1 --> E1
D1 --> E2
D1 --> F1
D1 --> F2
D1 --> F3
D1 --> F4
D1 --> F5
D1 --> F6
D1 --> F7
D1 --> F8
D2 --> E1
D2 --> E2
D2 --> F1
D2 --> F2
D2 --> F3
D2 --> F4
D2 --> F5
D2 --> F6
D2 --> F7
D2 --> F8
F4 --> G
F7 --> GPrimary Recommendation: A Rust-native, direct-to-metal cloud OS stack, optimized for performance, safety, and distributed consensus, with a modern SvelteKit operator UI.
| Category | Technology/Version | Rationale |
|---|---|---|
| Frontend | SvelteKit 2, TypeScript 5, Tailwind CSS 4 | Fast, modern UI, TypeScript for safety, Tailwind for utility-first design |
| Prompt Ops Layer | MCP (custom AI orchestration), AsyncAPI 3.0 | For AI-native, event-driven ops, natural language support |
| Backend/OS Services | Rust 1.77, Tokio 1.38, gRPC, Custom OS modules | Rust for memory safety/performance, Tokio for async IO, gRPC for module comms |
| Kubernetes Engine | Kubernetes 1.36.3 (docs) | Latest, stable, direct-to-metal, cgroup v2 mandatory as of 1.36 |
| Container Engine | Custom QCE (Rust, leveraging cgroups v2, namespaces), Podman 6 | Native container execution, Podman for registry, cgroup v2 for isolation |
| Distributed State | FoundationDB 7, etcd 3.6 | Strong consistency, distributed metadata, etcd for consensus, FoundationDB for metadata |
| Cache/Ephemeral | Redis 7 | Fast ephemeral state, caching |
| Networking | WireGuard, Cilium 1.16 (eBPF) | Secure programmable networking, eBPF-based firewalling |
| TLS/ACME | Cert-Manager | Automated certificate management |
| Observability | OpenTelemetry | Unified logging, metrics, tracing |
| CI/CD | GitHub Actions | Automated builds, tests, deploys |
| Supported Hardware | x86_64 (Intel, AMD), NVIDIA Ampere+, AMD MI, NVMe/SATA, 10/25/40/100GbE NICs | Direct-to-metal, broad enterprise support |
| Deployment | Self-hosted bare metal, Equinix Metal, AWS EC2 Bare Metal | Direct hardware access, hybrid support |
Technology Compatibility:
Table: nodes
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| node_type | ENUM | ('origin', 'link'), NOT NULL |
| hostname | VARCHAR(255) | NOT NULL, UNIQUE |
| ip_address | INET | NOT NULL, UNIQUE |
| status | ENUM | ('active', 'joining', 'offline', 'failed'), NOT NULL |
| hardware_info | JSONB | NOT NULL |
| created_at | TIMESTAMP | DEFAULT NOW() |
| updated_at | TIMESTAMP | DEFAULT NOW() |
| INDEX idx_nodes_type ON nodes(node_type) | ||
| INDEX idx_nodes_status ON nodes(status) |
Table: cluster_deployments
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| cluster_name | VARCHAR(255) | NOT NULL, UNIQUE |
| type | ENUM | ('kubernetes', 'container'), NOT NULL |
| status | ENUM | ('deploying', 'active', 'failed', 'deleting'), NOT NULL |
| node_pool | JSONB | NOT NULL |
| config | JSONB | NOT NULL |
| created_by | UUID | FK to users.id, NOT NULL |
| created_at | TIMESTAMP | DEFAULT NOW() |
| updated_at | TIMESTAMP | DEFAULT NOW() |
| INDEX idx_clusters_type ON cluster_deployments(type) | ||
| INDEX idx_clusters_status ON cluster_deployments(status) |
Table: node_cluster_membership
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| node_id | UUID | FK to nodes.id, NOT NULL |
| cluster_id | UUID | FK to cluster_deployments.id, NOT NULL |
| role | ENUM | ('master', 'worker'), NOT NULL |
| joined_at | TIMESTAMP | DEFAULT NOW() |
| UNIQUE(node_id, cluster_id) | ||
| INDEX idx_ncm_cluster_id ON node_cluster_membership(cluster_id) |
Table: containers
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| name | VARCHAR(255) | NOT NULL, UNIQUE |
| image | VARCHAR(512) | NOT NULL |
| node_id | UUID | FK to nodes.id, NOT NULL |
| status | ENUM | ('creating', 'running', 'stopped', 'failed'), NOT NULL |
| resources | JSONB | NOT NULL (cpu, gpu, memory, storage, net) |
| policy_id | UUID | FK to policies.id, NULLABLE |
| started_at | TIMESTAMP | |
| stopped_at | TIMESTAMP | |
| created_by | UUID | FK to users.id, NOT NULL |
| created_at | TIMESTAMP | DEFAULT NOW() |
| updated_at | TIMESTAMP | DEFAULT NOW() |
| INDEX idx_containers_status ON containers(status) | ||
| INDEX idx_containers_node_id ON containers(node_id) |
Table: consensus_state
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| key | VARCHAR(255) | NOT NULL, UNIQUE |
| value | JSONB | NOT NULL |
| version | BIGINT | NOT NULL |
| last_updated | TIMESTAMP | DEFAULT NOW() |
| INDEX idx_consensus_key ON consensus_state(key) |
Table: prompts
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| user_id | UUID | FK to users.id, NOT NULL |
| prompt_text | TEXT | NOT NULL |
| intent | VARCHAR(255) | NOT NULL |
| status | ENUM | ('received', 'processing', 'completed', 'error'), NOT NULL |
| result | JSONB | NULLABLE |
| error_message | TEXT | NULLABLE |
| created_at | TIMESTAMP | DEFAULT NOW() |
| updated_at | TIMESTAMP | DEFAULT NOW() |
| INDEX idx_prompts_status ON prompts(status) | ||
| INDEX idx_prompts_user_id ON prompts(user_id) |
Table: infra_services
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| service_type | ENUM | ('networking', 'firewall', 'dns', 'tls', 'load_balancer', 'storage', 'registry', 'access_control'), NOT NULL |
| config | JSONB | NOT NULL |
| policy_id | UUID | FK to policies.id, NULLABLE |
| node_id | UUID | FK to nodes.id, NULLABLE |
| created_at | TIMESTAMP | DEFAULT NOW() |
| updated_at | TIMESTAMP | DEFAULT NOW() |
| INDEX idx_infra_service_type ON infra_services(service_type) | ||
| INDEX idx_infra_service_node_id ON infra_services(node_id) |
Table: policies
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| name | VARCHAR(255) | NOT NULL, UNIQUE |
| type | ENUM | ('network', 'firewall', 'access', 'storage', 'container'), NOT NULL |
| definition | JSONB | NOT NULL |
| created_by | UUID | FK to users.id, NOT NULL |
| created_at | TIMESTAMP | DEFAULT NOW() |
| updated_at | TIMESTAMP | DEFAULT NOW() |
| INDEX idx_policies_type ON policies(type) |
Table: users
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| username | VARCHAR(255) | NOT NULL, UNIQUE |
| VARCHAR(255) | NOT NULL, UNIQUE | |
| password_hash | VARCHAR(255) | NOT NULL |
| role | ENUM | ('operator', 'admin', 'viewer'), NOT NULL |
| status | ENUM | ('active', 'disabled'), NOT NULL |
| created_at | TIMESTAMP | DEFAULT NOW() |
| updated_at | TIMESTAMP | DEFAULT NOW() |
| INDEX idx_users_role ON users(role) |
Table: hardware_inventory
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| node_id | UUID | FK to nodes.id, NOT NULL |
| cpu | JSONB | NOT NULL |
| gpu | JSONB | NULLABLE |
| storage | JSONB | NOT NULL |
| network | JSONB | NOT NULL |
| updated_at | TIMESTAMP | DEFAULT NOW() |
| INDEX idx_hardware_node_id ON hardware_inventory(node_id) |
Table: audit_logs
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| user_id | UUID | FK to users.id, NOT NULL |
| action | VARCHAR(255) | NOT NULL |
| target_type | VARCHAR(255) | NOT NULL |
| target_id | UUID | NULLABLE |
| timestamp | TIMESTAMP | DEFAULT NOW() |
| details | JSONB | NULLABLE |
| INDEX idx_audit_user_id ON audit_logs(user_id) | ||
| INDEX idx_audit_action ON audit_logs(action) |
Table: events
| Name | Type | Constraints |
|---|---|---|
| id | UUID | PK, DEFAULT gen_random_uuid() |
| event_type | VARCHAR(255) | NOT NULL |
| entity_type | VARCHAR(255) | NOT NULL |
| entity_id | UUID | NULLABLE |
| data | JSONB | NULLABLE |
| occurred_at | TIMESTAMP | DEFAULT NOW() |
| INDEX idx_events_type ON events(event_type) |
nodes ↔ hardware_inventory (1:1)nodes ↔ node_cluster_membership ↔ cluster_deployments (M:N)containers ↔ nodes (N:1)infra_services ↔ nodes (N:1, node_id nullable for global services)infra_services ↔ policies (N:1)erDiagram
nodes {
UUID id PK
ENUM node_type
VARCHAR hostname
INET ip_address
ENUM status
JSONB hardware_info
TIMESTAMP created_at
TIMESTAMP updated_at
}
cluster_deployments {
UUID id PK
VARCHAR cluster_name
ENUM type
ENUM status
JSONB node_pool
JSONB config
UUID created_by FK
TIMESTAMP created_at
TIMESTAMP updated_at
}
node_cluster_membership {
UUID id PK
UUID node_id FK
UUID cluster_id FK
ENUM role
TIMESTAMP joined_at
}
containers {
UUID id PK
VARCHAR name
VARCHAR image
UUID node_id FK
ENUM status
JSONB resources
UUID policy_id FK
TIMESTAMP started_at
TIMESTAMP stopped_at
UUID created_by FK
TIMESTAMP created_at
TIMESTAMP updated_at
}
consensus_state {
UUID id PK
VARCHAR key
JSONB value
BIGINT version
TIMESTAMP last_updated
}
prompts {
UUID id PK
UUID user_id FK
TEXT prompt_text
VARCHAR intent
ENUM status
JSONB result
TEXT error_message
TIMESTAMP created_at
TIMESTAMP updated_at
}
infra_services {
UUID id PK
ENUM service_type
JSONB config
UUID policy_id FK
UUID node_id FK
TIMESTAMP created_at
TIMESTAMP updated_at
}
policies {
UUID id PK
VARCHAR name
ENUM type
JSONB definition
UUID created_by FK
TIMESTAMP created_at
TIMESTAMP updated_at
}
users {
UUID id PK
VARCHAR username
VARCHAR email
VARCHAR password_hash
ENUM role
ENUM status
TIMESTAMP created_at
TIMESTAMP updated_at
}
hardware_inventory {
UUID id PK
UUID node_id FK
JSONB cpu
JSONB gpu
JSONB storage
JSONB network
TIMESTAMP updated_at
}
audit_logs {
UUID id PK
UUID user_id FK
VARCHAR action
VARCHAR target_type
UUID target_id
TIMESTAMP timestamp
JSONB details
}
events {
UUID id PK
VARCHAR event_type
VARCHAR entity_type
UUID entity_id
JSONB data
TIMESTAMP occurred_at
}
nodes ||--o{ hardware_inventory : has
nodes ||--o{ node_cluster_membership : member_of
cluster_deployments ||--o{ node_cluster_membership : contains
nodes ||--o{ containers : runs
users ||--o{ audit_logs : logs
users ||--o{ prompts : issues
users ||--o{ policies : defines
infra_services ||--o{ policies : enforces
infra_services ||--o{ nodes : runs_on
containers ||--o{ policies : governed_bynodes, cluster_deployments, and containers are partitioned by node location or physical cluster for scale-out./api/v1/GET /api/v1/nodesGET /api/v1/nodes/:idPOST /api/v1/nodes (register new LINK node)PATCH /api/v1/nodes/:id (update node status/config)GET /api/v1/clustersPOST /api/v1/clusters (create cluster)GET /api/v1/clusters/:idPATCH /api/v1/clusters/:idDELETE /api/v1/clusters/:idGET /api/v1/clusters/:id/nodesGET /api/v1/containersPOST /api/v1/containers (create/run container)GET /api/v1/containers/:idPATCH /api/v1/containers/:id (update resources/policy)POST /api/v1/containers/:id/stopPOST /api/v1/containers/:id/startDELETE /api/v1/containers/:idGET /api/v1/state/:keyPATCH /api/v1/state/:key (update state atomically)GET /api/v1/state/changes (stream state changes/events)POST /api/v1/prompts (submit prompt)GET /api/v1/prompts/:idGET /api/v1/prompts/:id/resultGET /api/v1/prompts/historyPOST /api/v1/prompts/:id/feedback (operator feedback/refinement)GET /api/v1/servicesGET /api/v1/services/:idPATCH /api/v1/services/:id (update config/policy)GET /api/v1/services/:id/policyPATCH /api/v1/services/:id/policyPOST /api/v1/services/:id/reloadGET /api/v1/policiesPOST /api/v1/policiesGET /api/v1/policies/:idPATCH /api/v1/policies/:idDELETE /api/v1/policies/:idGET /api/v1/usersPOST /api/v1/usersGET /api/v1/users/:idPATCH /api/v1/users/:idDELETE /api/v1/users/:idPOST /api/v1/auth/loginPOST /api/v1/auth/logoutPOST /api/v1/auth/refreshGET /api/v1/hardwareGET /api/v1/hardware/:node_idGET /api/v1/auditGET /api/v1/audit/:idGET /api/v1/eventsGET /api/v1/events/:idGET /api/v1/notificationsPOST /api/v1/notifications/acknowledge{
"name": "gpu-batch-1",
"image": "registry.example.com/ml:latest",
"node_id": "f7c2b9e0-3b8d-4dbe-8e57-7b2fa6d9e1e6",
"resources": {
"cpu": 8,
"gpu": 2,
"memory_gb": 64,
"storage_gb": 200,
"network_mbps": 10000
},
"policy_id": "d1a2c3e4-5678-90ab-cdef-1234567890ab"
}{
"id": "b6d5a2e1-4c3f-8b9d-7e2f-6d8c9b0a1e2f",
"status": "creating"
}flowchart TD
A["Operator UI / CLI"]
B["API Gateway (REST/AsyncAPI)"]
C1["Prompt Operations Service"]
C2["Cluster Orchestration Service"]
C3["Container Management Service"]
C4["Infra Services API"]
C5["Policy Management Service"]
C6["Access Control Service"]
C7["Hardware Inventory Service"]
C8["Audit/Event Service"]
A --> B
B --> C1
B --> C2
B --> C3
B --> C4
B --> C5
B --> C6
B --> C7
B --> C81. Direct-to-Metal Cloud OS Core
2. Kubernetes Engine (QKE)
3. Container Engine (QCE)
4. ORIGIN/LINK Node Consensus
5. AI-Native Prompt-Driven Ops Layer
6. Integrated Infrastructure Services
flowchart TD
A1["Prompt Interpreter"]
A2["Intent Validator"]
A3["Workflow Executor"]
A4["Feedback Engine"]
B1["Consensus Coordinator"]
B2["State Replicator"]
C1["QKE Orchestrator"]
C2["QCE Supervisor"]
D1["Networking Service"]
D2["Firewall Service"]
D3["DNS Service"]
D4["ACME/TLS Service"]
D5["Load Balancer Service"]
D6["Storage Service"]
D7["Container Registry"]
D8["Access Control Service"]
E["Distributed Metadata Store"]
A1 --> A2
A2 --> A3
A3 --> C1
A3 --> C2
A3 --> D1
A3 --> D2
A3 --> D3
A3 --> D4
A3 --> D5
A3 --> D6
A3 --> D7
A3 --> D8
C1 --> E
C2 --> E
D1 --> E
D2 --> E
D3 --> E
D4 --> E
D5 --> E
D6 --> E
D7 --> E
D8 --> E
B1 <--> B2
B1 <--> E
B2 <--> EsequenceDiagram
participant Operator
participant UI
participant PromptOps
participant IntentValidator
participant QKE_Orchestrator
participant ConsensusCoordinator
participant InfraServices
Operator->>UI: "Create Kubernetes cluster on 3 nodes"
UI->>PromptOps: Submit prompt
PromptOps->>IntentValidator: Parse & validate intent
IntentValidator->>ConsensusCoordinator: Check state/resources
ConsensusCoordinator-->>IntentValidator: State OK
IntentValidator-->>PromptOps: Intent valid
PromptOps->>QKE_Orchestrator: Provision cluster
QKE_Orchestrator->>InfraServices: Configure networking, storage, etc.
QKE_Orchestrator->>ConsensusCoordinator: Update cluster state
ConsensusCoordinator-->>QKE_Orchestrator: State confirmed
QKE_Orchestrator-->>PromptOps: Cluster provisioned
PromptOps-->>UI: Success feedback
UI-->>Operator: "Cluster created successfully"MCP (AI orchestration): Integrated within the prompt-driven ops layer. Uses AsyncAPI 3.0 events for intent exchange and feedback.
Protocol: AsyncAPI 3.0 over WebSocket/gRPC.
Auth: JWT for UI/CLI, mTLS for internal modules.
Data Format: JSON/YAML (AsyncAPI events).
Error Handling: All intent events include correlation IDs; errors returned as structured AsyncAPI error events with suggestions.
Cert-Manager module interacts with public ACME endpoints (e.g., Let's Encrypt).
Platform: Equinix Metal (single bare metal server for ORIGIN node) Containerization: Docker Compose for all user-space modules and the operator UI Essentials:
QBOS_NODE_TYPE=originQBOS_CLUSTER_NAME=mainJWT_SECRET, DATABASE_URL (FoundationDB/etcd), REDIS_URLDeferred to Phase 2+:
Estimated Setup Time: 4–8 hours (single operator, basic cloud CLI knowledge)
Deployment Topology Diagram (Phase 1)
flowchart TD
A["Operator Workstation"]
B["Operator UI (SvelteKit, Docker)"]
C["Prompt Ops Service (Docker)"]
D["QKE/QCE, Infra Services (Docker Compose)"]
E["FoundationDB/etcd (Docker)"]
F["Redis (Docker)"]
A --> B
B --> C
C --> D
D --> E
D --> FWhen to Reconsider:
VALIDATION CHECKLIST
End of Document
qbOS is a self-hosted, direct-to-metal cloud operating system that fundamentally reimagines cloud infrastructure management for enterprises, SaaS providers, research institutions, and edge computing vendors. By eliminating the traditional virtualization stack and integrating all critical infrastructure services (networking, firewall, DNS, ACME/TLS, load balancing, storage, registry, access control) into a single OS, qbOS delivers unmatched performance, operational simplicity, and security. Its unique ORIGIN and LINK distributed architecture, combined with an AI-native, prompt-driven operations layer, enables seamless, resilient, and intuitive management of cloud resources—provisioning production-ready Kubernetes clusters or containers on bare metal in under 60 seconds, all via natural language commands.
Key objectives:
Modern cloud infrastructure is plagued by complexity, inefficiency, and fragmentation. Enterprises and service providers must stitch together disparate services—often from multiple vendors—across layered virtualization stacks (hypervisors, VMs, containers, Kubernetes), leading to:
Current workarounds (e.g., VMware Cloud Foundation, OpenStack, DIY Kubernetes on bare metal, public cloud managed services) fail to deliver a truly integrated, direct-to-metal, AI-native experience. This is especially problematic for organizations with private/hybrid clouds, edge deployments, or on-premises infrastructure seeking bare-metal performance, simplified management, and strong security.
qbOS is a unified, direct-to-metal cloud OS that:
| Stakeholder | Role/Responsibility | Influence |
|---|---|---|
| Cloud Infrastructure Operators | Deploy, manage, and monitor qbOS environments; primary users of prompt-driven ops layer | High |
| DevOps Teams | Automate CI/CD, integrate qbOS into existing workflows | High |
| IT Administrators | Configure policies, enforce security/compliance | High |
| Enterprise Architects | Evaluate fit, ensure alignment with IT strategy | Medium |
| Security Officers | Audit, monitor, and enforce security/compliance | High |
| Executive Sponsors | Approve budget, strategic direction | High |
| Support/Consulting Partners | Provide deployment, integration, and training services | Medium |
| End Users (Developers, Data Scientists) | Consume provisioned clusters/resources, interact via self-service | Medium |
All requirements below are derived directly from the approved Architecture v1.0.0 and user Q&A. No deviations or generic placeholders.
| Category | Technology/Version |
|---|---|
| Frontend | SvelteKit 2, TypeScript 5, Tailwind CSS 4 |
| Prompt Ops | MCP (custom AI orchestration), AsyncAPI 3.0 |
| Backend/OS | Rust 1.77, Tokio 1.38, gRPC, Custom OS modules |
| Kubernetes | Kubernetes 1.36.3 (direct-to-metal, cgroup v2) |
| Container Engine | Custom QCE (Rust, cgroups v2, namespaces), Podman 6 |
| Distributed State | FoundationDB 7, etcd 3.6 |
| Cache | Redis 7 |
| Networking | WireGuard, Cilium 1.16 (eBPF) |
| TLS/ACME | Cert-Manager |
| Observability | OpenTelemetry |
| CI/CD | GitHub Actions |
| Hardware | x86_64 (Intel, AMD), NVIDIA Ampere+, AMD MI, NVMe/SATA, 10/25/40/100GbE NICs |
| Deployment | Self-hosted bare metal, Equinix Metal, AWS EC2 Bare Metal |
| Category | Requirement |
|---|---|
| Performance | Kubernetes clusters/containers provisioned in <60 seconds. Direct-to-metal execution. |
| Scalability | Horizontal: Add LINK nodes for scale-out; Vertical: Auto-detect and utilize hardware. |
| Reliability | Distributed consensus for state sync; automated failover; daily/weekly backups. |
| Availability | 99.99% uptime target for infra services; self-healing modules. |
| Security | Kernel-level isolation, eBPF firewall, RBAC, mTLS, JWT, encrypted data at rest/in transit. |
| Privacy | All user data encrypted (AES-256); audit logs immutable and time-stamped. |
| Compliance | Audit trails, policy versioning, GDPR-ready, supports data sovereignty. |
| Usability | Prompt-driven, natural language interface; accessible, responsive UI. |
| Observability | OpenTelemetry-based metrics, logs, traces; real-time dashboards. |
| Maintainability | Modular codebase, automated upgrades, rolling restarts, CI/CD pipelines. |
| Metric | Target/Goal | Measurement Method |
|---|---|---|
| Cluster Provisioning Time | < 60 seconds | Automated test, operator logs |
| Container Startup Time | < 10 seconds | Monitoring, logs |
| Operator Task Automation Rate | 70%+ of infra tasks via prompt ops | Usage analytics, operator surveys |
| System Uptime (infra services) | 99.99% | Monitoring, SLA reports |
| Security Incident Rate | < 1 per 12 months | Security monitoring, incident logs |
| Policy Compliance Audit Pass Rate | 100% | Audit logs, compliance reports |
| User Satisfaction (Operator NPS) | 70+ | Quarterly survey |
| Integration Time (to existing CI/CD) | < 1 day | Operator feedback, onboarding logs |
| Support Ticket Resolution Time | < 4 hours (P1), < 24 hours (P2) | Support system metrics |
| Risk Type | Description | Mitigation Strategy |
|---|---|---|
| Technical | Building robust, secure, performant direct-to-metal OS with integrated services | Phased rollout, extensive automated/unit/E2E testing, code audits |
| Market | Overcoming inertia of established stacks, operator retraining | Early design partner engagement, comprehensive training/docs |
| Competitive | Rapid response from incumbents, open-source alternatives | Emphasize unique AI-native, direct-to-metal, integrated approach |
| Execution | Seamless integration, security, and AI-native UX at scale | Modular architecture, CI/CD, regular user feedback loops |
| Security | New attack surfaces in direct-to-metal, AI-driven ops | Kernel hardening, eBPF, RBAC, regular CVE monitoring |
| Compliance | Meeting evolving regulatory requirements (GDPR, data sovereignty) | Policy-driven compliance, audit logs, regular legal review |
| Table | Purpose |
|---|---|
| nodes | ORIGIN/LINK node inventory and status |
| cluster_deployments | Kubernetes/container cluster lifecycle |
| node_cluster_membership | Node-to-cluster mapping |
| containers | Container lifecycle, resources, policy |
| consensus_state | Distributed state sync (key/value, versioned) |
| prompts | Prompt-driven ops, intent/result tracking |
| infra_services | Config/policy for networking, firewall, etc. |
| policies | Security/network/access/storage/container policies |
| users | Operator/admin/viewer accounts |
| hardware_inventory | CPU, GPU, storage, network per node |
| audit_logs | Immutable, time-stamped action logs |
| events | System/user events, for observability |
Refer to Architecture ER Diagram (see Architecture for full mermaid diagram).
API Style: REST/AsyncAPI for operator interfaces; gRPC for internal module comms
Versioning: /api/v1/
Authentication: JWT bearer tokens (operator APIs), mTLS (internal APIs)
Rate Limiting: 100 req/min per user (most endpoints), 20/min for destructive ops
GET /api/v1/nodesGET /api/v1/nodes/:idPOST /api/v1/nodes (register new LINK node)PATCH /api/v1/nodes/:id (update node status/config)GET /api/v1/clustersPOST /api/v1/clusters (create cluster)GET /api/v1/clusters/:idPATCH /api/v1/clusters/:idDELETE /api/v1/clusters/:idGET /api/v1/clusters/:id/nodes{
"name": "gpu-batch-1",
"image": "registry.example.com/ml:latest",
"node_id": "f7c2b9e0-3b8d-4dbe-8e57-7b2fa6d9e1e6",
"resources": {
"cpu": 8,
"gpu": 2,
"memory_gb": 64,
"storage_gb": 200,
"network_mbps": 10000
},
"policy_id": "d1a2c3e4-5678-90ab-cdef-1234567890ab"
}{
"id": "b6d5a2e1-4c3f-8b9d-7e2f-6d8c9b0a1e2f",
"status": "creating"
}| Phase | Deliverables | Timeline | Dependencies |
|---|---|---|---|
| Phase 1 | Single-node (ORIGIN) deployment, all core services, prompt-driven ops, API, UI, audit/logging | Month 1-3 | Hardware, Docker, CI/CD |
| Phase 2 | Multi-node (LINK), distributed consensus, auto-scaling, advanced observability, external integrations | Month 4-6 | Phase 1 completion, feedback |
| Phase 3 | Multi-region, CDN, service mesh, WAF, advanced compliance | Month 7-9 | Phase 2, regulatory review |
End of PRD
This PRD is fully aligned with the validated evaluation intelligence, approved Architecture v1.0.0, and all user-specified requirements. All sections are personalized for qbOS’s unique vision and target market. No generic content remains.
This epic covers critical non-functional requirements including system performance, scalability, security, monitoring, reliability, and usability to ensure qbOS meets enterprise-grade standards.
As an operator, I want the UI and prompt-driven operations layer to be accessible, responsive, and provide clear error handling, so that I can manage infrastructure efficiently and with minimal errors.
Priority: high
Dependencies: Implement Natural Language Command Parsing and Intent Mapping (31363ded-0fec-4105-92f5-22f66b3b921f), Develop Unified Management Interface for Infrastructure Services (46e1dd16-7ce2-418e-b60a-6609ffd93f1c)
Acceptance Criteria:
Story Points: 8
As the system, I want to achieve 99.99% uptime for infrastructure services and support automated backups and failover, so that qbOS is highly reliable and recoverable.
Priority: high
Dependencies: Implement ORIGIN Node Initialization and State Management (3c1aeba0-f8f5-4909-836a-865e20d69169), Ensure Resilience and Fault Tolerance in Distributed State Replication (2d245efa-8022-408f-a55f-af895de2fa70)
Acceptance Criteria:
Story Points: 8
As the system, I want to collect unified logs, metrics, and traces, and provide real-time alerts, so that operators can monitor system health and respond to incidents promptly.
Priority: high
Dependencies: Develop Unified Management Interface for Infrastructure Services (46e1dd16-7ce2-418e-b60a-6609ffd93f1c)
Acceptance Criteria:
Story Points: 8
As the system, I want to enforce strong authentication, authorization, and encryption mechanisms, so that qbOS meets enterprise security and compliance standards.
Priority: high
Dependencies: Implement Granular Policy Control and Enforcement for Infrastructure Services (1618e6c1-0c39-43d6-aa5b-c0e7758e5702)
Acceptance Criteria:
Story Points: 8
As the system, I want to support horizontal and vertical scaling to handle increasing numbers of concurrent users, nodes, and data volume, so that qbOS remains responsive and reliable at scale.
Priority: high
Dependencies: Implement LINK Node Dynamic Addition and State Synchronization (4da271ab-5e34-44c2-a04e-dc901fba35ef), Support High-Throughput Network Interfaces with Offloads (19d9780e-8d2b-4f8a-9d76-d99d5cf26d02)
Acceptance Criteria:
Story Points: 8
As the system, I want Kubernetes clusters and containers to be provisioned within target timeframes (<60 seconds for clusters, <10 seconds for containers), so that operators experience fast deployment.
Priority: high
Dependencies: Implement Bare Metal Kubernetes Cluster Provisioning (3973ff4e-9afa-4c9b-b3cd-b9a3a7086950), Implement Container Lifecycle Management on Direct Hardware (cc681990-182e-4c39-bf98-a6ce50d44adf)
Acceptance Criteria:
Story Points: 8
This epic covers the quality attributes ensuring that infrastructure services support fine-grained policy control and dynamic reconfiguration to meet security, performance, and compliance requirements.
As the system, I want to maintain consistency and correctness of policy enforcement across all infrastructure services, so that security and compliance requirements are reliably met.
Priority: high
Dependencies: Enable Dynamic Policy Definition and Modification by Operators (d2d11df1-f68b-4013-89d1-bf50ffde9fda)
Acceptance Criteria:
Story Points: 8
As an operator, I want to define and modify detailed policies dynamically via APIs and prompt-driven commands, so that I can adapt infrastructure configurations quickly to evolving operational needs.
Priority: high
Dependencies: Implement Granular Policy Control and Enforcement for Infrastructure Services (1618e6c1-0c39-43d6-aa5b-c0e7758e5702)
Acceptance Criteria:
Story Points: 8
This epic covers the implementation of security mechanisms enforcing container lifecycle management and resource isolation directly on hardware using OS-level primitives, integrated access control, and mandatory firewall policies.
As the system, I want to implement kernel-level isolation primitives and mandatory firewall policies for containers, so that container security is robust without virtualization overhead.
Priority: high
Dependencies: Enforce Secure Resource Isolation and Access Control for Containers (c430b8af-84b9-4fd6-866d-17db460f3272)
Acceptance Criteria:
Story Points: 8
As the system, I want to enforce strict CPU, GPU, memory, storage, and network resource allocation for containers, so that resource usage is predictable and isolated.
Priority: high
Dependencies: Implement Container Lifecycle Management on Direct Hardware (cc681990-182e-4c39-bf98-a6ce50d44adf)
Acceptance Criteria:
Story Points: 8
This epic covers the development and validation of a custom distributed consensus mechanism optimized for low-latency, peer-to-peer state synchronization without centralized controllers, ensuring unified state replication and conflict resolution.
As the system, I want the consensus protocol to be resilient and fault tolerant, so that the distributed cloud maintains availability and consistency despite node failures or network issues.
Priority: high
Dependencies: Design and Implement Low-Latency Consensus Protocol (9d001779-4371-4e25-a822-85cd90eeb785)
Acceptance Criteria:
Story Points: 8
As the system, I want to implement a distributed consensus protocol that provides low-latency synchronization between ORIGIN and LINK nodes, so that the distributed cloud state remains consistent and responsive.
Priority: high
Dependencies: Develop Custom Distributed Consensus Protocol for Low-Latency State Sync (67ade0f2-287a-4530-8287-8634f40667fb)
Acceptance Criteria:
Story Points: 13
This epic covers support for widely adopted enterprise-grade hardware including x86_64 CPUs, leading GPU architectures, enterprise NVMe and SATA storage, and high-throughput network interfaces with common offloads, ensuring qbOS runs reliably and efficiently.
As the system, I want to support 10/25/40/100 GbE network interfaces with offloads such as SR-IOV and RDMA, so that qbOS can provide high-performance networking for cloud workloads.
Priority: high
Dependencies: Integrate Essential Infrastructure Services into Unified System (92e9f606-cee9-4873-a31b-b8aaa6380e47)
Acceptance Criteria:
Story Points: 8
As the system, I want to support enterprise-grade NVMe and SATA storage devices, so that qbOS can provide reliable and performant storage services.
Priority: high
Dependencies: Integrate Essential Infrastructure Services into Unified System (92e9f606-cee9-4873-a31b-b8aaa6380e47)
Acceptance Criteria:
Story Points: 8
As the system, I want to support leading GPU architectures such as NVIDIA Ampere+ and AMD MI series, so that qbOS can run AI/ML workloads with full GPU access.
Priority: high
Dependencies: Support x86_64 CPUs with Hardware Virtualization Extensions (1a56b6e0-61f2-48d8-84a1-6d398a875093)
Acceptance Criteria:
Story Points: 8
As the system, I want to support x86_64 CPUs from major vendors with hardware virtualization extensions for future compatibility, so that qbOS can run efficiently on enterprise hardware.
Priority: high
Acceptance Criteria:
Story Points: 5
This epic covers the integration of MCP and AsyncAPI to enable natural language command interpretation, intent validation, and interactive feedback within the prompt-driven operations layer, enhancing usability and reducing operational complexity.
As the system, I want to use AsyncAPI to validate intents against system state and provide interactive feedback to operators, so that command execution is safe, predictable, and user-friendly.
Priority: high
Dependencies: Develop Intent Validation and Contextual Feedback Mechanism (59b771f1-cadc-4884-a13f-cc42d3e6c10d), Integrate MCP for AI-Orchestrated Natural Language Parsing (22b2d1e1-2e61-4da1-b63b-71170f5834ba)
Acceptance Criteria:
Story Points: 8
As the system, I want to integrate MCP into the prompt-driven operations layer to parse and interpret natural language commands, so that operators can interact with qbOS using AI-native commands.
Priority: high
Dependencies: Implement Natural Language Command Parsing and Intent Mapping (31363ded-0fec-4105-92f5-22f66b3b921f)
Acceptance Criteria:
Story Points: 8
This epic covers the development of integrated infrastructure services including networking, firewall, DNS, ACME/TLS, load balancing, storage, and access control, providing operators with granular policy control and dynamic configuration capabilities.
As an operator, I want to define and enforce detailed policies for security, networking, access control, and performance across all infrastructure services, so that I can ensure compliance and operational flexibility.
Priority: high
Dependencies: Develop Configuration Interfaces for All Integrated Infrastructure Services (f6bf1d96-d4b0-492b-8127-6d96b4185772), Enforce Secure Resource Isolation and Access Control for Containers (c430b8af-84b9-4fd6-866d-17db460f3272)
Acceptance Criteria:
Story Points: 13
As an operator, I want to configure networking, firewall, DNS, ACME/TLS, load balancing, storage, and access control services via APIs and prompt-driven commands, so that I can customize infrastructure to meet specific requirements.
Priority: high
Dependencies: Integrate Essential Infrastructure Services into Unified System (92e9f606-cee9-4873-a31b-b8aaa6380e47), Automate Execution of Validated Intents via Workflow Executor (a1f599a8-79f5-45ed-a62b-57bc8f8600f3)
Acceptance Criteria:
Story Points: 13
This epic covers the development of the natural-language driven operations interface leveraging MCP and AsyncAPI, enabling operators to manage infrastructure via natural language commands with intent interpretation, validation, and interactive feedback.
As the system, I want to automate the execution of validated intents through service APIs, so that complex infrastructure workflows are performed seamlessly and reliably.
Priority: high
Dependencies: Develop Intent Validation and Contextual Feedback Mechanism (59b771f1-cadc-4884-a13f-cc42d3e6c10d), Implement Bare Metal Kubernetes Cluster Provisioning (3973ff4e-9afa-4c9b-b3cd-b9a3a7086950), Implement Container Lifecycle Management on Direct Hardware (cc681990-182e-4c39-bf98-a6ce50d44adf), ()
Acceptance Criteria:
Story Points: 13
As the system, I want to validate parsed intents against current system state and policies, and provide clear, contextual feedback or suggestions for errors or ambiguities, so that operators can refine commands interactively and ensure safe operations.
Priority: high
Dependencies: Implement Natural Language Command Parsing and Intent Mapping (31363ded-0fec-4105-92f5-22f66b3b921f), Develop Unified Management Interface for Infrastructure Services (46e1dd16-7ce2-418e-b60a-6609ffd93f1c)
Acceptance Criteria:
Story Points: 8
As the system, I want to parse and understand natural language commands from operators, mapping them to actionable intents, so that operators can manage infrastructure easily using natural language.
Priority: high
Dependencies: Develop Unified Management Interface for Infrastructure Services (46e1dd16-7ce2-418e-b60a-6609ffd93f1c)
Acceptance Criteria:
Story Points: 8
This epic covers the implementation of the distributed cloud model using ORIGIN and LINK nodes that synchronize state directly via a custom distributed consensus mechanism optimized for low-latency, peer-to-peer synchronization without external control planes.
As the system, I want to implement a custom distributed consensus mechanism optimized for low-latency, peer-to-peer synchronization without centralized controllers, so that the distributed cloud maintains unified state replication and conflict resolution.
Priority: high
Dependencies: Implement ORIGIN Node Initialization and State Management (3c1aeba0-f8f5-4909-836a-865e20d69169), Implement LINK Node Dynamic Addition and State Synchronization (4da271ab-5e34-44c2-a04e-dc901fba35ef)
Acceptance Criteria:
Story Points: 13
As an operator, I want to add LINK nodes dynamically that synchronize state directly with ORIGIN and other LINK nodes, so that the distributed cloud scales seamlessly without external control planes.
Priority: high
Dependencies: Implement ORIGIN Node Initialization and State Management (3c1aeba0-f8f5-4909-836a-865e20d69169)
Acceptance Criteria:
Story Points: 13
As the system, I want the ORIGIN node to initialize and manage the distributed cloud state, so that it serves as the base for the distributed cloud.
Priority: high
Dependencies: Implement Direct Hardware Resource Management (bdd23bb5-2ae8-4e40-a7a7-c7d7ebbd8ae2)
Acceptance Criteria:
Story Points: 8
This epic covers the development of the Container Engine that runs containers directly on hardware with full CPU, GPU, storage, and network access, managing container lifecycle and resource isolation securely without virtualization overhead.
As the system, I want to enforce strict CPU, GPU, memory, storage, and network resource isolation and apply integrated access control and firewall policies to containers, so that container execution is secure and performant.
Priority: high
Dependencies: Implement Container Lifecycle Management on Direct Hardware (cc681990-182e-4c39-bf98-a6ce50d44adf), Integrate Essential Infrastructure Services into Unified System (92e9f606-cee9-4873-a31b-b8aaa6380e47)
Acceptance Criteria:
Story Points: 8
As an operator, I want to create, start, stop, and delete containers running directly on hardware, so that I can manage container workloads efficiently without virtualization layers.
Priority: high
Dependencies: Implement Direct Hardware Resource Management (bdd23bb5-2ae8-4e40-a7a7-c7d7ebbd8ae2), Integrate Essential Infrastructure Services into Unified System (92e9f606-cee9-4873-a31b-b8aaa6380e47)
Acceptance Criteria:
Story Points: 13
This epic covers the development of the Kubernetes Engine capable of provisioning production-ready Kubernetes clusters directly on bare metal hardware within 60 seconds, including cluster lifecycle management and scaling.
As an operator, I want to scale Kubernetes clusters and manage their lifecycle through the system, so that I can adapt cluster size and state to changing workload demands.
Priority: high
Dependencies: Implement Bare Metal Kubernetes Cluster Provisioning (3973ff4e-9afa-4c9b-b3cd-b9a3a7086950)
Acceptance Criteria:
Story Points: 8
As an operator, I want to deploy fully functional Kubernetes clusters directly on bare metal within 60 seconds, so that I can rapidly provision production-ready clusters without virtualization overhead.
Priority: high
Dependencies: Implement Direct Hardware Resource Management (bdd23bb5-2ae8-4e40-a7a7-c7d7ebbd8ae2), Integrate Essential Infrastructure Services into Unified System (92e9f606-cee9-4873-a31b-b8aaa6380e47)
Acceptance Criteria:
Story Points: 13
This epic covers the development of the core qbOS system that runs directly on hardware, replacing traditional virtualization stacks and integrating all essential cloud infrastructure services into a unified system for improved performance and simplified management.
As an operator, I want a unified interface to manage all integrated infrastructure services, so that I can simplify cloud infrastructure operations and reduce overhead.
Priority: high
Dependencies: Integrate Essential Infrastructure Services into Unified System (92e9f606-cee9-4873-a31b-b8aaa6380e47)
Acceptance Criteria:
Story Points: 8
As the system, I want to integrate networking, firewall, DNS, ACME/TLS, load balancing, storage, container registry, and access control services into one cohesive OS, so that operators can manage cloud infrastructure seamlessly.
Priority: high
Dependencies: Implement Direct Hardware Resource Management (bdd23bb5-2ae8-4e40-a7a7-c7d7ebbd8ae2)
Acceptance Criteria:
Story Points: 13
As the system, I want to manage hardware resources directly without hypervisor or VM layers, so that qbOS can provide high-performance direct-to-metal cloud operations.
Priority: high
Acceptance Criteria:
Story Points: 8
Spike
As a technical writer, I need to set up a documentation framework and create initial developer guides, so that the team and future contributors have clear references.
Priority: medium
Timebox: 4 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 5
Spike
As a developer, I want to set up code quality tools including linters and formatters, so that codebase maintains consistency and quality.
Priority: medium
Timebox: 2 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 3
Spike
As a QA engineer, I need to set up testing frameworks for unit, integration, and end-to-end tests, so that code quality and functionality are verified continuously.
Priority: high
Timebox: 4 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 8
Spike
As a backend engineer, I need to set up a database migration framework, so that schema changes can be managed and deployed safely.
Priority: high
Timebox: 3 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 5
Spike
As a DevOps engineer, I need to set up CI/CD pipelines using GitHub Actions, so that builds, tests, and deployments are automated and reliable.
Priority: high
Timebox: 5 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 8
Spike
As a developer, I need to set up local and Docker-based development environments, so that development and testing can be performed consistently across the team.
Priority: high
Timebox: 4 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 5
Spike
As a developer, I need to initialize the code repository and project scaffolding, so that the team has a standardized starting point for development.
Priority: high
Timebox: 3 days
Expected Outcomes:
Acceptance Criteria:
Story Points: 3
Execution order should follow dependencies; complete "Depends on" tasks before each task.
Acceptance Criteria:
Definition of Done:
81bd169e-1f73-4d56-b2c4-1eddfb71fad6Acceptance Criteria:
Definition of Done:
b85b076f-9082-485b-8d0e-c36b44d34380Acceptance Criteria:
Definition of Done:
b85b076f-9082-485b-8d0e-c36b44d34380Acceptance Criteria:
Definition of Done:
aeee9d2a-067d-4b59-9c47-6f6a2796708aAcceptance Criteria:
Definition of Done:
f0543287-e66e-4e5f-b698-c828f9935888, b85b076f-9082-485b-8d0e-c36b44d34380Acceptance Criteria:
Definition of Done:
b85b076f-9082-485b-8d0e-c36b44d34380, a3158e17-6592-40a7-b256-88ad3d538d9cAcceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
1c46fb65-1c91-46c3-956e-7190e7aa8fe2Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
81bd169e-1f73-4d56-b2c4-1eddfb71fad6Acceptance Criteria:
Definition of Done:
2b091d4e-8615-4607-a65e-78716784946dAcceptance Criteria:
Definition of Done:
8a8e605d-8245-4246-bc47-d3083ee55865Acceptance Criteria:
Definition of Done:
81bd169e-1f73-4d56-b2c4-1eddfb71fad6, 1c46fb65-1c91-46c3-956e-7190e7aa8fe2Acceptance Criteria:
Definition of Done:
f0543287-e66e-4e5f-b698-c828f9935888, 7265da3a-5ab1-4190-aeba-af519689aaa0Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
6f8f9117-9bb9-42d4-9078-43c354b2a43dAcceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
a93b8916-0977-452e-88ca-5eeff8ff8a03Acceptance Criteria:
Definition of Done:
8d94ffc9-095d-44a5-b822-03c5b3031821, b85b076f-9082-485b-8d0e-c36b44d34380Acceptance Criteria:
Definition of Done:
a93b8916-0977-452e-88ca-5eeff8ff8a03Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
8a8e605d-8245-4246-bc47-d3083ee55865, b6865ceb-327f-4c3a-a8ab-54904a92a0a1Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
a16d9bbd-b7a7-481a-b758-abceaa6083ebAcceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
a93b8916-0977-452e-88ca-5eeff8ff8a03Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
Acceptance Criteria:
Definition of Done:
users ↔ audit_logs (1:N)policies ↔ users (N:1)prompts ↔ users (N:1)POST /api/v1/clusters/:id/nodes (add node to cluster)DELETE /api/v1/clusters/:id/nodes/:node_idProtocol: HTTPS (ACME v2).
Auth: JWK-based client auth.
Retry: Exponential backoff on transient failures.
Container Registry (federated): OCI-compatible registry support (e.g., Docker Hub, Quay).
Protocol: HTTPS (OCI Registry API).
Auth: OAuth2/JWT tokens or registry-specific credentials.
Error Handling: Retries for network failures, fallback to local cache.
Hardware Drivers: All hardware integrations via OS kernel modules and user-space daemons (Rust, C).
POST /api/v1/clusters/:id/nodes (add node to cluster)DELETE /api/v1/clusters/:id/nodes/:node_idGET /api/v1/containersPOST /api/v1/containers (create/run container)GET /api/v1/containers/:idPATCH /api/v1/containers/:id (update resources/policy)POST /api/v1/containers/:id/stopPOST /api/v1/containers/:id/startDELETE /api/v1/containers/:idGET /api/v1/state/:keyPATCH /api/v1/state/:key (update state atomically)GET /api/v1/state/changes (stream state changes/events)POST /api/v1/prompts (submit prompt)GET /api/v1/prompts/:idGET /api/v1/prompts/:id/resultGET /api/v1/prompts/historyPOST /api/v1/prompts/:id/feedback (operator feedback/refinement)GET /api/v1/servicesGET /api/v1/services/:idPATCH /api/v1/services/:id (update config/policy)GET /api/v1/services/:id/policyPATCH /api/v1/services/:id/policyPOST /api/v1/services/:id/reloadGET /api/v1/policiesPOST /api/v1/policiesGET /api/v1/policies/:idPATCH /api/v1/policies/:idDELETE /api/v1/policies/:idGET /api/v1/usersPOST /api/v1/usersGET /api/v1/users/:idPATCH /api/v1/users/:idDELETE /api/v1/users/:idPOST /api/v1/auth/loginPOST /api/v1/auth/logoutPOST /api/v1/auth/refreshGET /api/v1/hardwareGET /api/v1/hardware/:node_idGET /api/v1/auditGET /api/v1/audit/:idGET /api/v1/eventsGET /api/v1/events/:idGET /api/v1/notificationsPOST /api/v1/notifications/acknowledgef6bf1d96-d4b0-492b-8127-6d96b4185772