BubbaBuilt

Self-hosted inference · OpenAI-compatible API

Your own LLM infrastructure, productised.

BubbaBuilt is the billing, metering and access layer in front of your GPU servers. Customers get a familiar API. You keep the hardware.

Your hardware, our control plane

Point the gateway at any OpenAI-compatible server — vLLM, Ollama, llama.cpp — and keep inference on your own GPUs.

Hashed API keys

Cryptographically random keys shown once, stored as hashes, with prefixes, rotation and per-key usage.

Rate limits & credits

Per-plan request limits and transactional credit accounting that can never go negative.

Usage metering

Every request recorded with tokens, latency, cost and charge — without storing prompts by default.

Tenant isolation

Row-level security on every customer-owned table, plus server-side authorization on every endpoint.

Model routing

Publish stable model names and swap the hardware behind them without breaking customer integrations.