Performance Tuning
Latency, throughput, and resource sizing for workloads
28 documentation pages carry this tag. All tags
- →
AI inference and RAG pipelines
- →
API rate limits
Reference
- →
Build a RAG Pipeline with Ollama and Qdrant
deployment
- →
CPU render-farm worker pool
Template
- →
Creator-AI inference worker
Template
- →
Deploy a Redis or Valkey cache with the redis-cache template
deployment
- →
Deploy the audio post-production worker template with OpenTofu
deployment
- →
Deploy the CPU render-farm worker pool template with OpenTofu
deployment
- →
Deploy the live RTMP/SRT ingest and restream template with OpenTofu
deployment
- →
Flavors
Explanation
- →
How to auto-scale a VM tier with Heat
How-to
- →
How to customize a template's image and flavor
How-to
- →
How to extend a block storage volume
How-to
- →
How to maintain a Quake AI VM over its lifetime
How-to
- →
How to put a CDN in front of a Quake AI workload
How-to
- →
How to resize a Quake AI VM to a dedicated-CPU plan
How-to
- →
How to use CloudFront with object storage buckets
How-to
- →
Indie multiplayer game servers
- →
Live RTMP/SRT ingest and restream
Template
- →
Live streaming and restream
- →
Media archive and VOD origin
- →
Model Resource Requirements
Reference
- →
Open lakehouse
- →
Redis / Valkey cache
Template
- →
Resource Tiers
Reference
- →
RPC and blockchain nodes
- →
Scientific computing and research
- →
Self-Hosted AI on Cloud Infrastructure
Explanation