x402-llm-gateway
activeAn x402 gateway for accessing large language model services.
Settled via Coinbase.
- Transactions · 30d
- 42
- Volume · 30d
- $1.44
- Unique buyers · 30d
- 4
- Uptime · 30d
- 64.7%
- Latency p50
- 222ms
- Reported calls · 30d
- 29
Endpoints (5 live)
POST/v1/vision— OpenAI-compatible vision (image understanding) from a private, self-hosted local vision-language model. Pay per call in USDC on Base via x402 - no API key or account. Send a public image URL plus a prompt (or OpenAI image_url messages); receive a standard chat.completion JSON answering about the image. Up to 4 images per call (more are clamped). Same privacy-first operator as our Private LLM Inference chat tiers. (0.02 USDC on Base)POST/v1/extended— Extended tier of a private, self-hosted LLM inference service: the highest output ceiling (32,768 tokens) on a local open-weight model for long-form writing, reports and full document drafts. OpenAI-compatible chat completions. Pay per call in USDC on Base via x402 - no API key or account. Pick the cheaper quick/long tiers for shorter work. (0.2 USDC on Base)POST/v1/long— Long-form tier of a private, self-hosted LLM inference service: an 8,192-token output ceiling on a local open-weight model for summarization, drafting and multi-paragraph answers. OpenAI-compatible chat completions: send a messages array, receive standard chat.completion JSON. Pay per call in USDC on Base via x402 - no API key or account. Use the cheaper quick tier (4,096 tokens) for short answers and tool calls. (0.05 USDC on Base)POST/v1/embeddings— Privacy-first OpenAI-compatible text embeddings (768 dims) from a local encoder on dedicated hardware - your text never reaches a third-party cloud. Send {"input": string or array of up to 64 texts}; receive standard OpenAI embeddings JSON with usage. Built for RAG, semantic search, dedupe and clustering in agent pipelines. ~2,000-token limit per text. Pay per call in USDC on Base (x402), no API key or account. (0.001 USDC on Base)POST/v1/chat/completions— Privacy-first chat API: prompts run on a local open-weight model on dedicated hardware - your data is never forwarded to OpenAI/Anthropic or any third-party LLM cloud. OpenAI-compatible chat completions: summarize, draft, extract, translate, code, answer questions. Send a messages array, get chat.completion JSON. Pay per call in USDC on Base (x402), no API key or account. Tiers by output cap: quick 4,096 / long 8,192 / extended 32,768 tokens. Free trial: POST /v1/trial. (0.005 USDC on Base)
First seen · last seen · last active