Skip to main content
← Transmute
Interactive Guide

How Transmute Works

An interactive walkthrough of the architecture, threat model, and audit trail that keeps your AI infrastructure secure, without changing how your team works.

What is Transmute?

Transmute is a transparent redaction gateway that sits between every AI tool your developers use and every LLM provider they call. It inspects every request and response in real time, scrubbing secrets, PII, and protected content before it ever leaves your environment.

Think of it as a security camera for your AI pipeline. It doesn't block your team, it watches, redacts, and logs. Developers keep using Copilot, Cursor, Claude Code, and their favorite CLI tools. Transmute runs silently in the background, catching what shouldn't leave the building.

No Call-Home

Runs on your infrastructure. No Purfect cloud, no AI-payload collection by Purfect Labs, the redaction mapping stays on your host. Sanitized requests still go to the model provider you configured.

Zero Code Changes

Set one environment variable. Every API call flows through Transmute automatically. No SDK required.

Provider Agnostic

Works with any OpenAI-compatible API. Anthropic, OpenAI, Google, local models. Transmute doesn't care.

Source Delivered

You own the code. No license server. No subscription. Full source at handoff.

Ready to lock down your AI pipeline?

One day to deploy. One env var to configure. Zero cloud dependencies. Source code is yours.

See Pricing Book a Demo

Frequently Asked Questions

Transmute operates as a transparent proxy. In local mode, added latency is typically under 5ms, imperceptible to users. The proxy sits on the same host, so there's no network hop. Filter pack matching uses pre-compiled regex trees that evaluate in microseconds.
Transmute is designed as a local sidecar, not a remote service. It runs on your infrastructure alongside your applications. If the Transmute process stops, it fails closed, requests are blocked rather than passed through uninspected. For high-availability deployments, you can run multiple Transmute instances behind a load balancer.
No TLS interception. Your AI client is configured (via a BASE_URL override) to send its requests to the Transmute gateway running on localhost, so Transmute receives the request in cleartext on your own machine. Transmute sanitizes it, replacing the secrets and PII its packs match with placeholders, then makes its own ordinary HTTPS call to the model provider you configured. The request still goes to the provider; Transmute removes the matched values first, and the redaction mapping stays on your host (and in opaque mode, isn't kept at all).
Transmute is provider-agnostic. It works with any OpenAI-compatible API (Anthropic, OpenAI, Google, DeepSeek, Mistral, local models via Ollama/vLLM). The proxy speaks standard HTTP/1.1 and doesn't care which model is on the other side.
WAFs and API gateways operate at the network/HTTP layer. They block IPs, rate-limit, and check headers. Transmute operates at the semantic layer: it understands the structure of LLM requests and responses. It can detect when a prompt contains secrets, when a model response leaks PII, or when a developer is about to send production configs to an external API. Traditional tools can't see inside JSON-RPC bodies.
No. Just run the Transmute app: it detects the AI tools on each machine and wires them up itself, so every API call flows through Transmute automatically. No environment variables, no SDK, no library, no code changes. Your developers keep using their existing tools and workflows.