LiliuxFlow

Your models. Your silicon. Your infrastructure.

A local LLM platform for Apple Silicon. Manage API access, model lifecycle and native inference in one place.

Apple SiliconOpenAI API
One workspace. A complete serving stack.Architecture illustration
LiliuxFlowControl plane
  1. ClientSDK · application
  2. LiteLLMAPI keys · routing · usage
  3. llama-swapModel lifecycle
  4. LilyNative inference
  5. MetalApple GPU
SSE streaming response · reasoning / content
Illustrative request and response flow. No live traffic or performance data.

From request to inference

Every layer, with a purpose.

An OpenAI-compatible API, backed by native inference.

  1. 01LiteLLM

    Control API access, enforce model permissions and track usage.

  2. 02llama-swap

    Route requests to models and manage the inference service lifecycle.

  3. 03Lily

    Run inference natively with Rust and Metal.

  4. 04Apple Silicon

    Run inference using the GPU and unified memory.

Designed for local operations

The essentials, integrated.

Manage models, API access and development workflows in one workspace.

01 /llama-swap

Model lifecycle, managed.

Manage model lifecycle and inference service settings in one console.

qwen3.8-flash-next-lily-q4-262k
Q4262,144 total contextSplit / MTP0
02 /LiteLLM

API access, under control.

Give each application its own virtual API key. Manage model permissions, budgets and usage.

Virtual API key•••• •••• ••••
03 /SSE

Responses, streamed.

Stream reasoning and content, with support for request cancellation.

04 /Apple Silicon

Unified memory, shared.

Run the model in Apple Silicon unified memory, with a total context limit of 262,144 tokens and resource safeguards.

05 /OpenAI API

Fits your workflow.

Use Chat Completions, model listings and your existing OpenAI SDK.

06 /Native stack

Native stack. Local execution.

Rust, Metal and native services. Manage models and services through the existing consoles.

An OpenAI-compatible API

Connect. Start building.

Set the base URL, add your virtual API key, then select a model.

API reference
POST /v1/chat/completionsGET /v1/models

Reasoning and tool calling depend on the model and runtime configuration. Validate tool arguments returned by the model before execution. See the API reference for limits.

from openai import OpenAI

client = OpenAI(
    base_url="https://liliuxflow.diurnoctra.com/v1",
    api_key="YOUR_VIRTUAL_KEY",
)

response = client.chat.completions.create(
    model="qwen3.8-flash-next-lily-q4-262k",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)
Connect with the OpenAI SDK/v1
Illustrative trace
  1. POST
  2. qwen3.8-flash-next-lily-q4-262k
  3. Lily
  4. Metal
  5. SSE

API endpoint example. Use your configured base URL and virtual API key.