Model lifecycle, managed.
Manage model lifecycle and inference service settings in one console.
qwen3.8-flash-next-lily-q4-262kYour models. Your silicon. Your infrastructure.
A local LLM platform for Apple Silicon. Manage API access, model lifecycle and native inference in one place.
From request to inference
An OpenAI-compatible API, backed by native inference.
Control API access, enforce model permissions and track usage.
Route requests to models and manage the inference service lifecycle.
Run inference natively with Rust and Metal.
Run inference using the GPU and unified memory.
Designed for local operations
Manage models, API access and development workflows in one workspace.
Manage model lifecycle and inference service settings in one console.
qwen3.8-flash-next-lily-q4-262kGive each application its own virtual API key. Manage model permissions, budgets and usage.
•••• •••• ••••Stream reasoning and content, with support for request cancellation.
Run the model in Apple Silicon unified memory, with a total context limit of 262,144 tokens and resource safeguards.
Use Chat Completions, model listings and your existing OpenAI SDK.
Rust, Metal and native services. Manage models and services through the existing consoles.
An OpenAI-compatible API
Set the base URL, add your virtual API key, then select a model.
API referencePOST /v1/chat/completionsGET /v1/modelsReasoning and tool calling depend on the model and runtime configuration. Validate tool arguments returned by the model before execution. See the API reference for limits.
from openai import OpenAI
client = OpenAI(
base_url="https://liliuxflow.diurnoctra.com/v1",
api_key="YOUR_VIRTUAL_KEY",
)
response = client.chat.completions.create(
model="qwen3.8-flash-next-lily-q4-262k",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)POSTqwen3.8-flash-next-lily-q4-262kLilyMetalSSEAPI endpoint example. Use your configured base URL and virtual API key.