NewGLM-5.3 is live as Abliterated Large v2.Try it now
ReferenceReviewed 2026-05-01

LLM image safety policies

Abliteration safety policies check image attachments before inference and record each decision.

Image safety checks run server-side on every image you send through the Abliteration gateway.

Locked platform protections and your effective organization and project policies are evaluated through Abliteration's self-hosted classifier.

Definition

LLM image safety policies

LLM image safety is a server-side policy check applied before an attached image reaches the model.

Why it matters
  • Block prohibited content from reaching the model regardless of which SDK the caller uses.
  • Return a stable policy error so client apps can handle blocked requests consistently.
  • Apply the same policy engine across /v1/chat/completions, /v1/messages, and /v1/responses.
  • Record policy decisions for audit and export workflows.
How it works
  1. 01Send a request with text + image content blocks as usual.
  2. 02The gateway evaluates the newest request against the locked platform baseline and effective customer policies.
  3. 03If an enforcing policy matches, the gateway returns HTTP 403 with code policy_violation.
  4. 04If everything passes, the request flows to vLLM/Modal for inference.
  5. 05Configure additional organization or project policies in the Console.
Rejection response shape
HTTP/1.1 403 Forbidden
Content-Type: application/json

{
  "error": {
    "message": "The request was blocked by a safety policy.",
    "type": "policy_error",
    "code": "policy_violation"
  }
}
FAQ

Frequently asked questions.

What categories are checked?

Abliteration's locked baseline protects against sexual content involving minors and self-harm content, intent, and instructions. Your organization can add plain-language policies for its own requirements.

How are multiple images handled?

The complete newest request is checked as one policy candidate. Any enforcing policy match rejects the request before model inference.

Does base64 vs HTTPS URL matter for moderation?

No. Both shapes pass through the same policy gate. Base64 data URLs skip the SSRF fetch but not safety evaluation.

Can I disable moderation?

Customers cannot disable Abliteration's locked baseline. Organization owners and admins can add, monitor, enforce, pause, and version their own policies.

What HTTP status and error code do rejections return?

HTTP 403 with error.code = 'policy_violation'. The decision is recorded in the organization's safety log.

Are moderation calls billed?

Abliteration does not add a separate safety-policy surcharge. Inference usage is billed through the normal model request.