PullO is a private AI network layer that turns local AI models into secure, shareable OpenAI-compatible APIs for your team. Connect Ollama, LM Studio, or any OpenAI-compatible endpoint. Expose local LLMs as production-ready APIs without port forwarding, tunnels, or cloud-hosted inference. Share AI endpoints securely with API keys, workspaces, and role-based access control. Supports streaming, MCP tool integration, rate limiting, and real-time analytics. Built for developer teams that need private AI infrastructure behind corporate firewalls.
Turn local AI models into team-accessible OpenAI-compatible APIs in four simple steps. Pull-based architecture keeps data on your machine.
Add the PullO Chrome extension (MV3). It acts as a secure WebSocket relay that pulls requests from the cloud queue to your local AI models.
PullO auto-discovers your local models on Ollama, LM Studio, or any OpenAI-compatible endpoint. No importing or migration required.
Create scoped API keys with per-model permissions, rate limits, and daily budgets. Share access without exposing your local infrastructure.
Point any OpenAI SDK, CLI tool, or pipeline at your PullO endpoint. Drop-in compatible with existing developer workflows and AI agents.
Everything you need to turn local AI models into a team-ready OpenAI-compatible API — no tunnels, no port forwarding, no cloud inference.
Pull-based architecture with zero inbound ports. Works behind corporate firewalls, NAT, and VPNs. No public IP or port forwarding needed. Data never leaves your machine — prompts and responses are never stored.
Drop-in replacement for OpenAI SDKs. Just change the base_url to your PullO endpoint. Works with any language or tool that speaks the OpenAI API format — Python, TypeScript, cURL, LangChain, and more.
Connect Ollama, LM Studio, or any OpenAI-compatible endpoint. Support for Llama, Mistral, Qwen, Gemma, DeepSeek, CodeGemma, Phi and all models available through your local runtime.
Persistent, low-latency WebSocket connection between the Chrome extension and cloud backend. Heartbeat health monitoring, streaming SSE responses, and automatic reconnection for reliable model serving.
Workspace isolation, role-based access (Owner, Admin, Member), API key management with SHA-256 hashing, per-model permissions, rate limits (RPM), and daily budget controls.
Built-in intelligence layer with web search, URL fetching, calculator, and date/time tools. Extend models with MCP (Model Context Protocol) servers and Corsair plugin marketplace for GitHub, Slack, and Gmail integrations.

Install the PullO Chrome extension, connect your local Ollama or LM Studio model, and start sharing private AI through a secure OpenAI-compatible API endpoint.
Connect Ollama & LM Studio directly in browser
Tell us what's not working, request new features, or let us know how PullO can better serve your local AI workflow.