Skip to content
Solutions

Self-hosted vibecode stack

Self-hosted vibecode stack

The modern vibecode stack is assembled from a sprawl of managed services, each separately metered: hosting on one vendor, Postgres on another, Redis on a third, then vector search, inference, auth, object storage, queues, and observability, each with its own meter and invoice. This page maps each of those categories to a self-host path on Quake AI, so you can run the stack on Compute you operate.

What this is for#

Teams that have watched a vibecode project accumulate a dozen separately billed SaaS dependencies can consolidate those dependencies onto one stack they own. Each category below has a validated OpenTofu template and a deployment walkthrough that stands up the self-hosted equivalent on Quake AI Compute.

The mechanism this addresses is fragmentation. A typical vibecode stack runs:

  • A separate bill, dashboard, and contract for each vendor, reconciled at different times.
  • Independent scaling cliffs, where each vendor moves from free tier to paid and paid to "contact sales" on its own curve.
  • Metered data movement between layers, because traffic between the hosting tier, the database, the cache, and the inference provider crosses billed egress boundaries on every hop.

Running the stack on Quake AI puts those layers on one provider you operate, with no metered outbound transfer. You take on operating the services yourself in exchange for one bill and flat egress between layers. The trade-off is real: you run, patch, back up, and scale each service under the shared responsibility model.

Reference architecture#

DeveloperPaid SaaS sprawlQuake AI, one stack you ownHostingDB, cache, vectorsInferenceAnalytics, telemetry, metricsPM, chat, secretsApp runtimeData layerAI layerIdentityDev platformObservabilityTeam opsNext.js or container appPostgres + pgvectorRedisQdrantObject StorageInference gatewayKeycloakForgejo + CIHarborUmamiGlitchTipMonitoring stackUptime KumaPlaneMattermostInfisicalUnleash many meters, many billsone stack, flat egress
Click to zoom
A vibecode stack consolidated onto Quake AI: the app runtime, data layer, AI layer, developer platform, observability, and team operations all run on Compute you operate, replacing a sprawl of separately metered SaaS

Download diagram: SVG, PNG, and PDF.

The stack groups into six layers, each assembled from templates you can deploy independently:

  • App runtime runs the long-lived application and its API.
  • Data layer holds the relational database, cache, vector store, and object storage.
  • AI layer routes inference requests to a model backend.
  • Dev platform hosts the git forge, CI, and container registry.
  • Observability covers product analytics, error telemetry, infrastructure metrics, and uptime.
  • Team operations covers project tracking, chat, secrets, feature flags, and the other tools a team pays for separately from the app itself.

Map your stack to a self-host path#

Each row names a common paid SaaS for the category and the Quake AI template or deployment that replaces it. Build the stack one row at a time; nothing here requires adopting every layer at once.

CategoryCommon paid SaaSSelf-host path on Quake AI
App hosting and PaaSVercel, Netlify, RenderCoolify host, Next.js app, containerized app
Relational databaseNeon, Amazon RDSSelf-managed Postgres
MySQL-compatible databasePlanetScale, Amazon Aurora MySQLMySQL/MariaDB database
Cache and key-valueUpstash, ElastiCacheRedis
Vector searchPinecone, Weaviate CloudQdrant; Postgres carries the pgvector extension for smaller corpora
Full backend (BaaS)Supabase, FirebaseSupabase self-host stack bundles Postgres, auth, storage, and Realtime on one instance; or compose Postgres, Object Storage, and an auth provider (next row) yourself
Auth and identityClerk, Auth0Self-hosted OIDC identity provider (Keycloak); the Outline deployment shows an OIDC wiring end to end
Object storageAmazon S3, Cloudflare R2Object Storage with the bucket ACL template
Queues and background jobsInngest, QStash, SQSRedis as a broker; n8n for scheduled and event-driven workflows
Inference and model routingReplicate, Together, ModalInference gateway; see AI inference and RAG
Workflow automationZapier, Maken8n
Git and CIGitHub, GitLabForgejo with a bundled Actions runner
Container registryDocker Hub, GitHub Container RegistryHarbor
Knowledge base and docsNotion, ConfluenceOutline
Product and web analyticsPostHog, Mixpanel, Vercel AnalyticsUmami
Error telemetrySentry, LogRocketGlitchTip speaks the Sentry protocol; official Sentry SDKs report to it unchanged
Infrastructure metricsDatadog, Grafana CloudMonitoring stack
Uptime and status pagesStatuspage, PingdomUptime Kuma

Perimeter and gateways#

The rows above cover app runtime, data, AI, and team-ops SaaS. The perimeter layer (reverse proxy, WAF, API gateway, tunnel access, cache, edge functions, DNS) is a separate set of vendors. Self-host the software on regional VMs where a validated OpenTofu template exists; use a third-party CDN for global PoPs and volumetric DDoS absorption.

CategoryCommon paid SaaSSelf-host path on Quake AI
Reverse proxy + TLSCloudflare proxy, Fastly edge TLSEdge reverse proxy (deploy walkthrough)
WAF (self-managed L7)Cloudflare WAF, AWS WAF, Google Cloud ArmorEdge WAF appliance (deploy walkthrough); or a CDN-layer WAF via Front with a WAF
API gatewayCloudflare API Shield, AWS API Gateway, ApigeeAPI gateway (deploy walkthrough)
Tunnel + accessCloudflare Tunnel, Cloudflare AccessEdge tunnel gateway (deploy walkthrough)
HTTP cache (regional)Cloudflare cache, Fastly cacheEdge cache; geographic caching via Front with a CDN
Edge functionsCloudflare Workers, Vercel Edge, Deno DeployEdge functions (deploy walkthrough); Wasm on Kubernetes via Kubernetes cluster; Deno via Supabase self-host
Authoritative DNSRoute 53, Cloudflare DNSSelf-host authoritative DNS (PowerDNS on a Compute VM you operate; CoreDNS for in-cluster service discovery)
Global CDN + volumetric DDoSCloudflare CDN (global PoPs, L3/L4 scrubbing)Not replaced by self-hosting on Quake AI. Front with a CDN for geographic edge and provider-managed volumetric protection

The Edge & gateways category holds perimeter templates: edge reverse proxy, edge WAF, API gateway, edge tunnel gateway, edge cache, and edge functions. Each pairs with a deployment walkthrough linked in the table above where one exists.

Run the team's tools too#

The stack sprawl is not only technical. Teams that consolidate hosting and datastores often still meter project tracking, chat, and internal tools across a separate set of vendors. The same self-host path applies.

CategoryCommon paid SaaSSelf-host path on Quake AI
Secrets managementDoppler, 1Password SecretsInfisical; see Secrets management
Project and work managementLinear, JiraPlane; see Project and work management
Feature flags and experimentsLaunchDarkly, FlagsmithUnleash; see Feature flags and experiments
Internal tools and admin panelsRetoolAppsmith; see Internal tools and admin panels
Testing and CI browser gridsBrowserStack, Sauce LabsSelenium Grid; see Testing and QA infrastructure
Team chatSlackMattermost; see Team chat
Scheduling and bookingCalendlyCal.com; see Scheduling and booking
Files and collaborationDropbox, Google WorkspaceNextcloud; see Files and collaboration
Whiteboarding and diagrammingFigma, MiroExcalidraw; see Whiteboarding and diagramming

Services involved#

ServiceRole in this architectureDocs
ComputeHosts the app runtime, datastores, AI gateway, dev platform, observability, and team-ops instancesCompute
NetworkPrivate networks, security groups, and floating IPs for each instanceNetwork
Object StorageS3-compatible store for uploads, backups, and build artifactsObject Storage
Kubernetes (Magnum)Optional orchestration target for the layers that scale horizontallyKubernetes

Get started#

Start with the layer that costs you the most today, then add others. Each template links its deployment walkthrough.

Estimate the cost#

The estimate below covers a starting point: the application runtime and its database. Every other row from the tables above appears below as an optional add-on (its marginal cost on top of the starting point) or an alternative (its own standalone cost, to compare against the starting point instead of adding to it).

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on a mix of shared and dedicated vCPU.

Starting template$150.70/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

App tier

s1a.small · 2 shared vCPU, 2 GiB RAM, 0.5 Gbps

$16.50/mo

PostgreSQL database

m2a.xlarge · 4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

s1a.small

2 shared vCPU, 2 GiB RAM, 0.5 Gbps

$16.50

m2a.xlarge

4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00

Compute + RAM rate basis

6 vCPU + 18 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (90 GiB)

90 GiB at $0.08/GiB/mo

$7.20

Public IP (included)

1 included with the custom package

$0.00

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Configure your estimate

Check the add-ons you plan to deploy to build a monthly total. Nothing is selected to start, so the total below begins at the baseline.

Starting template

The required baseline, always included.

$150.70/mo
Your configured estimate$150.70/mo

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$51.70/mo

Production on the configured CPU

The headline estimate above; predictable steady-load performance.

$150.70/mo

Saves $99.00/mo while you build on shared CPU.

Shared flavors carry less RAM (m2a.xlarge (16 GiB RAM) -> s1a.medium (4 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Considerations and limits#

  • You operate every layer. Quake AI provides Compute, Network, and Object Storage; the services that run on them are yours under the shared responsibility model. There is no platform-managed database, cache, or inference product.
  • Single-node starting templates. Most templates here provision one VM with file-backed or block storage. For high availability, run each service against external Postgres and Object Storage and route requests through an edge reverse proxy or API gateway.
  • CPU-only compute. All instances are AMD EPYC with no GPU option (compute FAQ). The inference gateway routes to a model backend you run elsewhere or to a CPU-served model; it does not add local GPU.
  • Full BaaS. The Supabase self-host stack reproduces Supabase's database, auth, storage, and Realtime layer on one instance. For a backend assembled from separate pieces instead, compose Postgres, Object Storage, and the auth-oidc identity provider yourself.
  • Flat egress between layers. Quake AI applies a no-egress-fee policy, so traffic between the runtime, the database, and the cache does not accrue per-gigabyte charges the way inter-vendor traffic does.
  • Three US regions. All current regions are in the United States.
  • When this pattern is not the right fit.
    • Teams that want managed products with no servers to operate and accept per-vendor metering for that.
    • Workloads that need a GPU-backed managed inference endpoint with no instance to run, where a hosted inference provider behind the gateway fits better.
    • Early prototypes where a single free tier is enough and operating any infrastructure is premature.
Was this page helpful?