For developers
Open models. Lower prices. Two lines of code.
An OpenAI-compatible API for open models. Change the base URL and your key; keep your code.
Python
from openai import OpenAI client = OpenAI(base_url= "https://api.ai.pearlsafe.xyz/v1") chat = client.chat.completionsstream = chat.create(model=MODEL, messages=messages, stream=True)Open models, priced per token.
| Model | Input / 1M | Output / 1M | |
|---|---|---|---|
| Llama 3.1 8B Instruct | 32K | $0.017 | $0.034 |
| Gemma 4 31B IT | 32K | $0.119 | $0.34 |
The first request after a quiet period can take a couple of minutes while a GPU starts.
Keep your code; change two lines.
Base URL: https://api.ai.pearlsafe.xyz/v1
from openai import OpenAI
client = OpenAI(base_url="https://api.ai.pearlsafe.xyz/v1", api_key="YOUR_API_KEY")
stream = client.chat.completions.create(
model="llama-3.1-8b-instruct",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")-
Streaming
Server-sent events, like OpenAI.
-
Zero retention
We don't store prompts or answers, and never train on them.
-
Pay as you go
Per token, below the cheapest listed price for the same model.Per token, for the tokens you use.
Standard requests run on GPUs that also mine: our pool briefly checks a small slice of intermediate model values in memory, then discards them. Private requests are never mined. Compute privacy