PullO — Private AI Network Layer for Local Models

PullO is a private AI network layer that turns local AI models into secure, shareable OpenAI-compatible APIs for your team. Connect Ollama, LM Studio, or any OpenAI-compatible endpoint. Expose local LLMs as production-ready APIs without port forwarding, tunnels, or cloud-hosted inference. Share AI endpoints securely with API keys, workspaces, and role-based access control. Supports streaming, MCP tool integration, rate limiting, and real-time analytics. Built for developer teams that need private AI infrastructure behind corporate firewalls.

Workflow

How PullO Works

Turn local AI models into team-accessible OpenAI-compatible APIs in four simple steps. Pull-based architecture keeps data on your machine.

1. Install Extension

Add the PullO Chrome extension (MV3). It acts as a secure WebSocket relay that pulls requests from the cloud queue to your local AI models.

2. Connect Ollama

PullO auto-discovers your local models on Ollama, LM Studio, or any OpenAI-compatible endpoint. No importing or migration required.

3. Generate API Keys

Create scoped API keys with per-model permissions, rate limits, and daily budgets. Share access without exposing your local infrastructure.

4. Integrate Anywhere

Point any OpenAI SDK, CLI tool, or pipeline at your PullO endpoint. Drop-in compatible with existing developer workflows and AI agents.

Everything included

Private AI infrastructure for your team

Everything you need to turn local AI models into a team-ready OpenAI-compatible API — no tunnels, no port forwarding, no cloud inference.

Private AI Network Layer

Pull-based architecture with zero inbound ports. Works behind corporate firewalls, NAT, and VPNs. No public IP or port forwarding needed. Data never leaves your machine — prompts and responses are never stored.

OpenAI-Compatible API

Drop-in replacement for OpenAI SDKs. Just change the base_url to your PullO endpoint. Works with any language or tool that speaks the OpenAI API format — Python, TypeScript, cURL, LangChain, and more.

Any Local Model

Connect Ollama, LM Studio, or any OpenAI-compatible endpoint. Support for Llama, Mistral, Qwen, Gemma, DeepSeek, CodeGemma, Phi and all models available through your local runtime.

WebSocket Relay Architecture

Persistent, low-latency WebSocket connection between the Chrome extension and cloud backend. Heartbeat health monitoring, streaming SSE responses, and automatic reconnection for reliable model serving.

Team Access Controls

Workspace isolation, role-based access (Owner, Admin, Member), API key management with SHA-256 hashing, per-model permissions, rate limits (RPM), and daily budget controls.

MCP & Tool Integration

Built-in intelligence layer with web search, URL fetching, calculator, and date/time tools. Extend models with MCP (Model Context Protocol) servers and Corsair plugin marketplace for GitHub, Slack, and Gmail integrations.

0ms
Network latency
100%
On-device
40+
Models supported
$0
Per-token cost

Go from local model to team API in minutes.

Install the PullO Chrome extension, connect your local Ollama or LM Studio model, and start sharing private AI through a secure OpenAI-compatible API endpoint.

Chrome Extension

Connect Ollama & LM Studio directly in browser

Unpacked developer build · Free & Open Source · 2-minute setup guide
Install PullO Extension
Chrome MV3 · WebSocket relay
Install Ollama
Local inference server
Download a Model
Llama · Mistral · Qwen · Gemma · DeepSeek
Start Browsing
Step 4

Have a Query? Connect with us

Tell us what's not working, request new features, or let us know how PullO can better serve your local AI workflow.

0/1000