Self-hosted vibecode stack
Self-hosted vibecode stack
The modern vibecode stack is assembled from a sprawl of managed services, each separately metered: hosting on one vendor, Postgres on another, Redis on a third, then vector search, inference, auth, object storage, queues, and observability, each with its own meter and invoice. This page maps each of those categories to a self-host path on Quake AI, so you can run the stack on Compute you operate.
What this is for#
Teams that have watched a vibecode project accumulate a dozen separately billed SaaS dependencies can consolidate those dependencies onto one stack they own. Each category below has a validated OpenTofu template and a deployment walkthrough that stands up the self-hosted equivalent on Quake AI Compute.
The mechanism this addresses is fragmentation. A typical vibecode stack runs:
- A separate bill, dashboard, and contract for each vendor, reconciled at different times.
- Independent scaling cliffs, where each vendor moves from free tier to paid and paid to "contact sales" on its own curve.
- Metered data movement between layers, because traffic between the hosting tier, the database, the cache, and the inference provider crosses billed egress boundaries on every hop.
Running the stack on Quake AI puts those layers on one provider you operate, with no metered outbound transfer. You take on operating the services yourself in exchange for one bill and flat egress between layers. The trade-off is real: you run, patch, back up, and scale each service under the shared responsibility model.
Reference architecture#
Download diagram: SVG, PNG, and PDF.
The stack groups into six layers, each assembled from templates you can deploy independently:
- App runtime runs the long-lived application and its API.
- Data layer holds the relational database, cache, vector store, and object storage.
- AI layer routes inference requests to a model backend.
- Dev platform hosts the git forge, CI, and container registry.
- Observability covers product analytics, error telemetry, infrastructure metrics, and uptime.
- Team operations covers project tracking, chat, secrets, feature flags, and the other tools a team pays for separately from the app itself.
Map your stack to a self-host path#
Each row names a common paid SaaS for the category and the Quake AI template or deployment that replaces it. Build the stack one row at a time; nothing here requires adopting every layer at once.
| Category | Common paid SaaS | Self-host path on Quake AI |
|---|---|---|
| App hosting and PaaS | Vercel, Netlify, Render | Coolify host, Next.js app, containerized app |
| Relational database | Neon, Amazon RDS | Self-managed Postgres |
| MySQL-compatible database | PlanetScale, Amazon Aurora MySQL | MySQL/MariaDB database |
| Cache and key-value | Upstash, ElastiCache | Redis |
| Vector search | Pinecone, Weaviate Cloud | Qdrant; Postgres carries the pgvector extension for smaller corpora |
| Full backend (BaaS) | Supabase, Firebase | Supabase self-host stack bundles Postgres, auth, storage, and Realtime on one instance; or compose Postgres, Object Storage, and an auth provider (next row) yourself |
| Auth and identity | Clerk, Auth0 | Self-hosted OIDC identity provider (Keycloak); the Outline deployment shows an OIDC wiring end to end |
| Object storage | Amazon S3, Cloudflare R2 | Object Storage with the bucket ACL template |
| Queues and background jobs | Inngest, QStash, SQS | Redis as a broker; n8n for scheduled and event-driven workflows |
| Inference and model routing | Replicate, Together, Modal | Inference gateway; see AI inference and RAG |
| Workflow automation | Zapier, Make | n8n |
| Git and CI | GitHub, GitLab | Forgejo with a bundled Actions runner |
| Container registry | Docker Hub, GitHub Container Registry | Harbor |
| Knowledge base and docs | Notion, Confluence | Outline |
| Product and web analytics | PostHog, Mixpanel, Vercel Analytics | Umami |
| Error telemetry | Sentry, LogRocket | GlitchTip speaks the Sentry protocol; official Sentry SDKs report to it unchanged |
| Infrastructure metrics | Datadog, Grafana Cloud | Monitoring stack |
| Uptime and status pages | Statuspage, Pingdom | Uptime Kuma |
Perimeter and gateways#
The rows above cover app runtime, data, AI, and team-ops SaaS. The perimeter layer (reverse proxy, WAF, API gateway, tunnel access, cache, edge functions, DNS) is a separate set of vendors. Self-host the software on regional VMs where a validated OpenTofu template exists; use a third-party CDN for global PoPs and volumetric DDoS absorption.
| Category | Common paid SaaS | Self-host path on Quake AI |
|---|---|---|
| Reverse proxy + TLS | Cloudflare proxy, Fastly edge TLS | Edge reverse proxy (deploy walkthrough) |
| WAF (self-managed L7) | Cloudflare WAF, AWS WAF, Google Cloud Armor | Edge WAF appliance (deploy walkthrough); or a CDN-layer WAF via Front with a WAF |
| API gateway | Cloudflare API Shield, AWS API Gateway, Apigee | API gateway (deploy walkthrough) |
| Tunnel + access | Cloudflare Tunnel, Cloudflare Access | Edge tunnel gateway (deploy walkthrough) |
| HTTP cache (regional) | Cloudflare cache, Fastly cache | Edge cache; geographic caching via Front with a CDN |
| Edge functions | Cloudflare Workers, Vercel Edge, Deno Deploy | Edge functions (deploy walkthrough); Wasm on Kubernetes via Kubernetes cluster; Deno via Supabase self-host |
| Authoritative DNS | Route 53, Cloudflare DNS | Self-host authoritative DNS (PowerDNS on a Compute VM you operate; CoreDNS for in-cluster service discovery) |
| Global CDN + volumetric DDoS | Cloudflare CDN (global PoPs, L3/L4 scrubbing) | Not replaced by self-hosting on Quake AI. Front with a CDN for geographic edge and provider-managed volumetric protection |
The Edge & gateways category holds perimeter templates: edge reverse proxy, edge WAF, API gateway, edge tunnel gateway, edge cache, and edge functions. Each pairs with a deployment walkthrough linked in the table above where one exists.
Run the team's tools too#
The stack sprawl is not only technical. Teams that consolidate hosting and datastores often still meter project tracking, chat, and internal tools across a separate set of vendors. The same self-host path applies.
| Category | Common paid SaaS | Self-host path on Quake AI |
|---|---|---|
| Secrets management | Doppler, 1Password Secrets | Infisical; see Secrets management |
| Project and work management | Linear, Jira | Plane; see Project and work management |
| Feature flags and experiments | LaunchDarkly, Flagsmith | Unleash; see Feature flags and experiments |
| Internal tools and admin panels | Retool | Appsmith; see Internal tools and admin panels |
| Testing and CI browser grids | BrowserStack, Sauce Labs | Selenium Grid; see Testing and QA infrastructure |
| Team chat | Slack | Mattermost; see Team chat |
| Scheduling and booking | Calendly | Cal.com; see Scheduling and booking |
| Files and collaboration | Dropbox, Google Workspace | Nextcloud; see Files and collaboration |
| Whiteboarding and diagramming | Figma, Miro | Excalidraw; see Whiteboarding and diagramming |
Services involved#
| Service | Role in this architecture | Docs |
|---|---|---|
| Compute | Hosts the app runtime, datastores, AI gateway, dev platform, observability, and team-ops instances | Compute |
| Network | Private networks, security groups, and floating IPs for each instance | Network |
| Object Storage | S3-compatible store for uploads, backups, and build artifacts | Object Storage |
| Kubernetes (Magnum) | Optional orchestration target for the layers that scale horizontally | Kubernetes |
Get started#
Start with the layer that costs you the most today, then add others. Each template links its deployment walkthrough.
- Runtime first. Next.js app and its deploy walkthrough, or Coolify host for push-to-deploy across several apps.
- Add the data layer. Self-managed Postgres, Redis, and Qdrant for vector search.
- Add the developer platform. Forgejo with CI and Harbor, threaded end to end in CI/CD pipelines.
- Add observability. Umami for analytics, the monitoring stack for metrics, and Uptime Kuma for status pages.
- Add inference. Inference gateway, covered in AI inference and RAG.
- Add team operations. Plane for project tracking, Mattermost for chat, and Infisical for secrets.
Estimate the cost#
The estimate below covers a starting point: the application runtime and its database. Every other row from the tables above appears below as an optional add-on (its marginal cost on top of the starting point) or an alternative (its own standalone cost, to compare against the starting point instead of adding to it).
Monthly cost estimate
Pricing calculator ↗Sized as a custom package on a mix of shared and dedicated vCPU.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
App tier
s1a.small · 2 shared vCPU, 2 GiB RAM, 0.5 Gbps
PostgreSQL database
m2a.xlarge · 4 dedicated vCPU, 16 GiB RAM, 1 Gbps
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Included in baseline
s1a.small
2 shared vCPU, 2 GiB RAM, 0.5 Gbps
m2a.xlarge
4 dedicated vCPU, 16 GiB RAM, 1 Gbps
Compute + RAM rate basis
6 vCPU + 18 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.
Block storage (90 GiB)
90 GiB at $0.08/GiB/mo
Public IP (included)
1 included with the custom package
Package promotional discount
Flat −$5.00/mo on the custom package (same promotion as named plans).
Included at no charge
These line items are zero on Quake AI. Many other providers meter them separately.
Data transfer (inbound and outbound)
Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.
AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.
Learn morePrivate networking
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
Configure your estimate
Check the add-ons you plan to deploy to build a monthly total. Nothing is selected to start, so the total below begins at the baseline.
Starting template
The required baseline, always included.
Dev/test vs production
Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.
Dev/test on shared CPU
Burstable s1a flavors; suited to prototyping and low or bursty load.
Production on the configured CPU
The headline estimate above; predictable steady-load performance.
Saves $99.00/mo while you build on shared CPU.
Shared flavors carry less RAM (m2a.xlarge (16 GiB RAM) -> s1a.medium (4 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.
Pricing data last validated: . For current rates, check quake.ai/pricing.
Considerations and limits#
- You operate every layer. Quake AI provides Compute, Network, and Object Storage; the services that run on them are yours under the shared responsibility model. There is no platform-managed database, cache, or inference product.
- Single-node starting templates. Most templates here provision one VM with file-backed or block storage. For high availability, run each service against external Postgres and Object Storage and route requests through an edge reverse proxy or API gateway.
- CPU-only compute. All instances are AMD EPYC with no GPU option (compute FAQ). The inference gateway routes to a model backend you run elsewhere or to a CPU-served model; it does not add local GPU.
- Full BaaS. The Supabase self-host stack reproduces Supabase's database, auth, storage, and Realtime layer on one instance. For a backend assembled from separate pieces instead, compose Postgres, Object Storage, and the auth-oidc identity provider yourself.
- Flat egress between layers. Quake AI applies a no-egress-fee policy, so traffic between the runtime, the database, and the cache does not accrue per-gigabyte charges the way inter-vendor traffic does.
- Three US regions. All current regions are in the United States.
- When this pattern is not the right fit.
- Teams that want managed products with no servers to operate and accept per-vendor metering for that.
- Workloads that need a GPU-backed managed inference endpoint with no instance to run, where a hosted inference provider behind the gateway fits better.
- Early prototypes where a single free tier is enough and operating any infrastructure is premature.