CoderAI is a multimodal, multi-backend local model orchestrator with a drop-in OpenAI-compatible API server, so you can run models on your own GPUs instead of someone else's cloud.
It supports NVIDIA (CUDA), AMD and Intel (Vulkan) backends with per-model configuration, on-demand load/unload, smart VRAMβRAMβdisk offloading, prompt caching and aggregation, parallel execution and a request queue.
Each model also picks its own inference engine: Transformers, llama.cpp, ds4, colibri, kimi-k3-in-c, ktransformers/SGLang or vLLM β so a frontier-size mixture-of-experts model can stream off disk on hardware that could never hold it. A model can even be served on a GPU rented by the second through RunPod, with price and spend caps, idle teardown and a reaper that terminates pods nobody is using.
A full Web Studio covers chat, image, video and audio generation β text-to-image/video, image-to-video, upscaling, inpainting, face-swap, TTS/STT, voice cloning, dubbing and custom multi-step pipelines β alongside document OCR with schema-driven extraction, speaker recognition, multi-vector embeddings with reranking, and on-GPU LoRA training.
Screenshots




