NewGLM-5.3 is live as Abliterated Large v2.Try it now
ReferenceReviewed 2026-09-20

Multimodal LLM API (text + images)

OpenAI-compatible multimodal API accepting text and images across chat, Responses, and Messages endpoints.

A multimodal LLM API accepts text and images in the same request.

abliteration.ai exposes a multimodal API that mirrors OpenAI Chat Completions, so existing SDKs work without changes beyond switching the base URL.

Definition

Multimodal LLM API (text + images)

A multimodal LLM API accepts text and images in one request and returns natural-language or structured outputs grounded in both inputs.

Why it matters
  • Describe screenshots, diagrams, or scanned documents alongside text instructions.
  • Compare multiple images in one request.
  • Extract structured data such as tables, captions, and counts from images.
  • Reduce round-trips by sending all the context the model needs in one request.
How it works
  1. 01Send a single /v1/chat/completions request whose messages array contains user-role messages with multipart content.
  2. 02Each message.content array mixes { type: 'text', text } and { type: 'image_url', image_url: { url } } parts.
  3. 03Use HTTPS URLs or data: URLs. The backend validates and normalizes images before inference.
  4. 04Authenticate with a JWT or API key. Anonymous free-tier callers can attach images.
  5. 05Stream responses with stream: true; delta chunks contain text-only output regardless of input type.
Mixed text and image request
curl https://api.abliteration.ai/v1/chat/completions \
  -H "Authorization: Bearer $ABLIT_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "abliterated-model",
    "messages": [
      {
        "role": "user",
        "content": [
          { "type": "text", "text": "Compare these images — what changed?" },
          { "type": "image_url", "image_url": { "url": "https://example.com/before.jpg" } },
          { "type": "image_url", "image_url": { "url": "https://example.com/after.jpg" } }
        ]
      }
    ]
  }'
FAQ

Frequently asked questions.

What input types does the API accept?

Use text and image_url content parts on Chat Completions. Mix them in any order inside a single message.content array.

Which endpoints support which inputs?

/v1/chat/completions, /v1/messages, /v1/responses, and their /policy/... variants accept text and image inputs for abliterated-model.

What are the size limits?

15 MB per image for PNG, JPEG, WEBP, and GIF. The total request body cap is 35 MB after base64 encoding.

Can I send multiple images in the same request?

Yes. The effective image count depends on your plan and project settings. Latency and token use increase with each image.

Is the API OpenAI-compatible?

Yes. Use the OpenAI Python or Node SDK with baseURL pointing at https://api.abliteration.ai/v1. Existing chat-completions code works unmodified for text and images.

Does anon free-tier work?

Yes. Anonymous callers send X-Free-Tier: true and get one free request that may include supported image inputs.

What about streaming?

Set stream: true. SSE delta chunks come back the same way regardless of input type. Time to first token is higher for image inputs because the model processes the images before generating.