Skip to main content

Overview

The local LLM agent runs a quantized language model entirely in-process using Transformers.js (ONNX/WebAssembly). No external server, no base URL, and no API key are required. The model is downloaded from HuggingFace Hub on first run and cached locally.

Project Structure

Skills

Agent

app/agent.ts
The model (onnx-community/Qwen2.5-1.5B-Instruct, q4 quantized) is loaded via @browser-ai/transformers-js, which is an official Vercel AI SDK community provider for Transformers.js.

Payment

This agent is free (scheme: "free") — no x402 payment is required to call it.

Running locally

Docker deployment

Because the local LLM model (~1 GB of weights) is downloaded and prewarmed at image build time, the container starts serving requests immediately with no cold-start delay.

Build the image

Or use the npm script:

Run the container

Or:
Endpoints available at http://localhost:3000: