# NudgeBee — Full Reference > NudgeBee is an open source, self-hosted agentic AI platform for day-2 cloud operations. It replaces tool-sprawl and manual toil by turning a flood of disconnected alerts into ranked, correlated incidents, investigating each to a defensible evidence-cited root cause, surfacing every dollar of cloud and Kubernetes waste with the exact fix attached, and remediating manually, on a schedule, or autonomously with approval — all without changing the observability, cloud, ticketing, or identity stack the team already runs. Four AI assistants (AI-SRE, AI-FinOps, AI-K8s Ops, and CloudOps) run on one shared backend and are delivered through a single assistant named Nubi. ## About NudgeBee NudgeBee is a self-hosted, agentic CloudOps platform: AI-SRE, AI-FinOps, AI-K8s Ops, and a no-code Agentic Automation Builder on one backend. It runs inside the customer's own Kubernetes cluster and watches Kubernetes plus AWS, Azure, and GCP accounts, turning raw signals into ranked findings and walking operators through investigation and remediation. - **Who it is for:** SRE, Platform Engineering, DevOps, Cloud, and FinOps teams running Kubernetes and AWS/Azure/GCP. - **The assistant is branded "Nubi"** and is fully white-labelable (name, company, logo, colors) for MSP and OEM partners. - **Deployment:** self-hosted in the buyer's own cluster (VPC) or Cloud SaaS. - **Elevator pitch:** one AI copilot that investigates incidents to a real root cause, finds and fixes cloud and Kubernetes waste, and automates the fixes — inside your own cluster, no data leaves, no model lock-in. ## The four assistants - **AI-SRE** — autonomous incident investigation to an evidence-cited 5-Whys root cause. The AI is instructed to reject symptom-only answers: an HTTP 404/500/503, a CrashLoopBackOff, or a "Resource Missing" is explicitly not accepted as a root cause; a cited chain of evidence is required. Root-cause findings can be written back to incident tools. - **AI-FinOps** — continuous multi-cloud and Kubernetes cost optimization with dollar-quantified, ranked, one-click-apply recommendations across right-sizing, security, infrastructure upgrades, and configuration. Folds in first-party cloud savings signals and Kubernetes right-sizing, ranked by a FinOps score. - **AI-K8s Ops** — live Kubernetes topology, owner-level event de-noising (crash-loops, image-pull backoff, OOM kills, and more collapse to one issue per workload), ML anomaly detection, and in-browser pod exec, logs, metrics, and traces. - **Agentic Automation Builder** — no-code, durable, crash-safe runbooks triggered by chat, alert, or schedule; auto-pilot remediation with dry-run and guardrails. ## Key differentiators - **Investigate, not execute:** the AI runs read-only diagnostics freely; every create, update, or delete is classified and gated behind explicit human approval. It cannot change infrastructure on its own — enforced both in prompts and in the tool-access layer. - **No telemetry / no phone-home:** nothing leaves the cluster except integrations the operator wires up. - **Data stays in the cluster:** metrics, logs, and traces are queried in place through the customer's own observability stack; embeddings can run fully on-device for air-gapped environments. - **Outbound-only connection model:** the in-cluster agent dials out over one secure channel; zero inbound ports, no exposed kube-apiserver, no VPN. - **No model lock-in (bring your own model):** works across AWS Bedrock, OpenAI, Azure OpenAI, Google AI (Gemini), Vertex AI, SageMaker, HuggingFace, and Anthropic, plus on-device embeddings. - **Knowledge-graph correlation:** a real topology graph drives causal incident correlation — distinguishing root cause from symptom — instead of just time-window grouping. - **Multi-cloud parity:** AWS, Azure, GCP, and Kubernetes in one collector, one data model, and one recommendation taxonomy. - **Deeply white-labelable** from a single theme file (UI, emails, and bot), with no rebuild — built for MSP and OEM partners. - **SOC2 certified.** ## Integrations (overview) - **Clouds & platforms:** AWS, Azure, GCP, Cloud Foundry, Kubernetes (via an in-cluster agent). - **Observability, APM, logging & tracing:** Prometheus/Alertmanager, VictoriaMetrics, Datadog, Dynatrace, New Relic, Splunk Observability, Chronosphere, Last9, Observe, SolarWinds, Loggly, Azure App Insights, Elasticsearch, Loki, SigNoz, ClickHouse, OpenTelemetry, Jaeger, Grafana Tempo, Grafana, and eBPF. - **Ticketing / ITSM / incident:** Jira, ServiceNow, PagerDuty, Zenduty, GitHub Issues, GitLab Issues. - **ChatOps:** Slack, Microsoft Teams, Google Chat, Discord, and email — conversational (mention Nubi in-channel), not slash-command. - **Source control / GitOps:** GitHub, GitLab, Bitbucket, ArgoCD, Confluence. - **Identity / SSO:** Google, Okta, Azure AD, OneLogin, LDAP/AD, and email magic-link. - **LLM / AI providers:** Bedrock (default), OpenAI, Azure OpenAI, Google AI, Vertex AI, SageMaker, HuggingFace, Anthropic; on-device embeddings. ## Product pages - [Home](https://nudgebee.com/): Platform overview — the four AI assistants and the Agentic Automation Builder, self-hosted with no model lock-in and human-in-the-loop control. - [Pricing](https://nudgebee.com/pricing): Scope-based pricing. Pick AI-SRE, AI-FinOps, or both; self-hosted or Cloud SaaS; bring your own model; free tier available. - [Book a Demo](https://nudgebee.com/demo): A live product walkthrough of AI-powered incident management, cloud cost optimization, and workflow automation, tailored to your stack (typically 30–45 minutes). - [Contact](https://nudgebee.com/contact): Reach the NudgeBee team for demos, pricing, partnerships, or technical support at contact@nudgebee.com. ## Case studies - [Case studies overview](https://nudgebee.com/case-studies): Customer outcomes across AI-SRE, AI-FinOps, and AI-K8s Ops. - [E-commerce incident automation](https://nudgebee.com/case-studies/ecommerce-incident-automation): A top-10 e-commerce platform used agentic SRE workflows to automate incident triage, end $180K/hr outages, and resolve 61% of incidents without human intervention. - [Fintech cloud cost](https://nudgebee.com/case-studies/fintech-cloud-cost): A payments platform used the AI-FinOps assistant to cut cloud spend 34% and recover $2.4M in waste across 8 AWS accounts, with no FinOps headcount. - [Healthcare enterprise cloud savings](https://nudgebee.com/case-studies/healthcare-enterprise): A healthcare enterprise saved $1.2M annually and reduced cloud waste by 40% with zero additional headcount. - [SaaS Kubernetes optimization](https://nudgebee.com/case-studies/saas-kubernetes-optimization): A B2B SaaS company right-sized 200+ Kubernetes clusters, saved $1.2M/year, and reduced cluster management from 3 days to 15 minutes per month. ## Blog: incident response, SRE & MTTR - [What Is Root Cause Analysis in Engineering? Complete RCA Guide](https://nudgebee.com/resources/blog/root-cause-analysis-in-engineering): What RCA is, common methods, and how to run it in engineering teams. - [7 Best AI Tools for Root Cause Analysis in 2026](https://nudgebee.com/resources/blog/ai-tools-for-root-cause-analysis): A comparison of AI-assisted RCA tools for faster, more accurate diagnosis. - [AI Alert Investigation: How AI Speeds Up Incident Response](https://nudgebee.com/resources/blog/ai-alert-investigation): How AI triages and investigates alerts to cut time-to-diagnosis. - [Automated Incident Management: Benefits, Workflows & Examples](https://nudgebee.com/resources/blog/automated-incident-management): The benefits, workflows, and real examples of automating incident management. - [7 Best Automated Incident Response Tools in 2026](https://nudgebee.com/resources/blog/7-best-automated-incident-response-tools-in-2026): A round-up of the leading automated incident response tools. - [Best AI SRE Tools for Reducing MTTR](https://nudgebee.com/resources/blog/best-ai-sre-tools-for-reducing-mttr): AI SRE tools that measurably reduce mean time to resolution. - [Top 5 AI SRE Tools in 2026](https://nudgebee.com/resources/blog/best-sre-platforms-2025): The top AI SRE platforms compared for 2026. - [7 Best AI Tools for Reliability Engineers in 2026](https://nudgebee.com/resources/blog/best-ai-tools-for-reliability-engineers): Tooling that helps reliability engineers work faster and more reliably. - [Top 7 Incident Management Software for Enterprise in 2026](https://nudgebee.com/resources/blog/best-incident-management-software-for-enterprise-in-2026): Enterprise incident management platforms compared. - [7 Best Practices for Incident Management in Large Enterprises](https://nudgebee.com/resources/blog/best-practices-for-incident-management-in-large-enterprises): Practical incident-management practices for large organizations. - [Incident Management vs Problem Management: What's the Difference?](https://nudgebee.com/resources/blog/incident-management-vs-problem-management): How incident management and problem management differ and overlap. - [How to Reduce MTTR for Higher Reliability](https://nudgebee.com/resources/blog/how-to-reduce-mttr-proven-strategies-for-faster-recovery-and-higher-reliability): Proven strategies for faster recovery and higher reliability. - [What Is MTTR: ROI of Reducing MTTR Using Automated Response Systems](https://nudgebee.com/resources/blog/what-is-mttr-roi-of-reducing-mttr): What MTTR is and the ROI of reducing it with automation. - [7 Proven Ways to Reduce Incident Response Time in 2026](https://nudgebee.com/resources/blog/ways-to-reduce-incident-response-time): Concrete ways to shorten incident response time. - [7 Best Tools for Faster DevOps Incident Recovery (Compared)](https://nudgebee.com/resources/blog/tools-for-faster-devops-incident-recovery): DevOps incident-recovery tools compared. - [8 Best Open-Source Incident Investigation Tools for DevOps Teams](https://nudgebee.com/resources/blog/open-source-incident-investigation-tools-for-devops): Open-source options for investigating incidents. - [SRE Reliability & Observability Explained: Metrics, Logs & Traces](https://nudgebee.com/resources/blog/sre-reliability-and-observability): The fundamentals of reliability and observability for SRE teams. - [How AI Helps Improve Code Reliability in Modern Software Engineering](https://nudgebee.com/resources/blog/how-ai-helps-with-code-reliability): Ways AI improves code reliability across the software lifecycle. ## Blog: FinOps & cloud cost - [AI FinOps Agents: 9 Best Tools Compared (2026)](https://nudgebee.com/resources/blog/ai-finops-agents): A comparison of AI FinOps agents for cloud cost optimization. - [Top Cloud Automation Tools to Streamline Cloud Optimization in 2026](https://nudgebee.com/resources/blog/top-cloud-automation-tools-to-streamline-cloud-optimization-in-2025): Cloud automation tools that streamline cost and operations. ## Blog: Kubernetes troubleshooting - [How to Fix Kubernetes 502 Bad Gateway Error (Complete Guide)](https://nudgebee.com/resources/blog/fix-kubernetes-502-bad-gateway-error): Causes and fixes for Kubernetes 502 Bad Gateway errors. - [Exit Code 137 Kubernetes: Causes & Fix (OOMKilled Pods)](https://nudgebee.com/resources/blog/fixing-exit-code-137-pod-termination-kubernetes): Why pods get OOMKilled with exit code 137 and how to fix it. - [Readiness Probe Failed in Kubernetes: How to Fix It](https://nudgebee.com/resources/blog/readiness-probe-failed-in-kubernetes): Diagnosing and resolving failed readiness probes. - [Kubernetes Node Not Ready? How to Fix It Fast](https://nudgebee.com/resources/blog/troubleshoot-kubernetes-node-not-ready-error): Fast fixes for the Kubernetes "Node Not Ready" state. ## Blog: AI architecture & engineering - [KG vs RAG: Why SRE Teams Need Both](https://nudgebee.com/resources/blog/kg-vs-rag): Why knowledge graphs and retrieval-augmented generation complement each other for SRE. - [Knowledge Graphs vs Vector Databases for Enterprise AI](https://nudgebee.com/resources/blog/the-quiet-defeat-of-vector-databases): Trade-offs between knowledge graphs and vector databases for enterprise AI. - [Enterprise Context Layer: The Hidden Cost of Scaling AI](https://nudgebee.com/resources/blog/the-enterprise-context-layer): Why a shared context layer matters when scaling enterprise AI. - [Build vs Buy in AIOps: 10 Reasons DIY AI Agents Fail](https://nudgebee.com/resources/blog/the-building-trap-in-aiops): The pitfalls of building your own AIOps agents versus buying. - [7 Best AIOps Platforms for Startups and Enterprises in 2026](https://nudgebee.com/resources/blog/best-aiops-platforms-for-startups-and-enterprises-in-2025): AIOps platforms compared for startups and enterprises. - [CLI vs MCP: What We Use Where, and Why](https://nudgebee.com/resources/blog/cli-vs-mcp-at-nudgebee): How NudgeBee chooses between CLI tools and MCP for its agents. - [Building LLM Applications: RAG, Agents & Tool Calling](https://nudgebee.com/resources/blog/building-llm-applications): A practical guide to RAG, agents, and tool calling in LLM apps. - [Anatomy of an LLM Request: Prefill, Decode, Latency & Cost](https://nudgebee.com/resources/blog/anatomy-of-an-llm-request): How an LLM request works and what drives latency and cost. - [LLM Serving Cheat Sheet: GPUs, vLLM Flags & Cost Levers](https://nudgebee.com/resources/blog/llm-serving-cheat-sheet): A reference for serving LLMs efficiently on GPUs. - [The Future of DevOps: AI, Automation & Platform Engineering](https://nudgebee.com/resources/blog/future-of-devops-in-2026-and-beyond): Where DevOps is heading with AI, automation, and platform engineering. ## Blog: comparisons & alternatives - [NudgeBee vs PagerDuty AIOps: Which Platform Is Better?](https://nudgebee.com/resources/blog/nudgebee-vs-pagerduty-aiops): A head-to-head comparison of NudgeBee and PagerDuty AIOps. - [7 Best PagerDuty Competitors and Alternatives in 2026](https://nudgebee.com/resources/blog/pagerduty-competitors-and-alternatives): Leading alternatives to PagerDuty. - [7 Best Resolve AI Alternatives for SRE Teams in 2026](https://nudgebee.com/resources/blog/resolve-ai-alternatives): Alternatives to Resolve AI for SRE teams. ## Company - [NudgeBee Raises $3M Seed Led by Kalaari Capital](https://nudgebee.com/resources/blog/nudgebee-raises-3m-seed-led-by-kalaari-capital): NudgeBee's $3M seed round led by Kalaari Capital. - [Blog home](https://nudgebee.com/resources/blog): All NudgeBee articles on AI-SRE, FinOps, Kubernetes, and cloud operations. ## Optional - [Privacy Policy](https://nudgebee.com/privacy-policy): How NudgeBee collects, uses, and protects personal information. - [Terms of Service](https://nudgebee.com/terms): Terms for using the NudgeBee website and platform.