You can redact sensitive data from a document by asking an AI agent in plain language, and the redaction runs locally, inside the file, on your machine. With GroupDocs.Redaction.Mcp connected to Claude Desktop, Claude Code, Cursor or GitHub Copilot, one prompt is enough:

Redact every email address in case-file.pdf and tell me which file you saved.

The step-by-step version with config and troubleshooting is in the documentation: How to redact sensitive data from documents with AI agents.

Why can’t a chat model just redact the document for you?

A language model that reads a document and writes a new one is producing a rewrite, not a redaction. The rewrite can miss an occurrence, change wording you did not ask it to change, and it never touches the original file. A black rectangle drawn in a PDF viewer has a similar flaw: the text under the box is still in the file and can be copied out. The agent should not do the removing. It should propose the pattern and call a tool, and the engine should replace the matched text in the document itself.

That is the division of labour in GroupDocs.Redaction.Mcp. The agent proposes patterns and drives the calls. The engine applies them exactly. There is no classifier deciding what looks sensitive, so the behaviour is reviewable, and the decision about what counts as sensitive stays with you.

Way 1: Replace text by pattern

Use this for anything you can describe as a regular expression: email addresses, national IDs, phone numbers, account numbers, a client name. The agent calls redact_text, which replaces text matching the pattern in the document and saves a redacted copy. The original is untouched.

Redact anything matching \b\d{3}-\d{2}-\d{4}\b, US social security numbers, in case-file.pdf. Show me the pattern before you run it.

Asking to see the pattern first is worth the extra turn. A regex you have read is a redaction you can defend, and a loose one removes more than you meant while a tight one leaves data behind. The replacement parameter defaults to [REDACTED], which keeps the output easy to review.

Expected result: a redacted copy in the output folder, and an agent message naming the saved file.

Way 2: Clear the places a text pattern cannot reach

Cleaning the body text and shipping a file whose margin comments and author field still name the person is the most common way a redaction fails. Sensitive data sits in four places, and each has its own tool:

Where the data is Tool
Body text redact_text
Scans, photos, signatures redact_image_area
Comments and sticky notes redact_annotations
Author, company, dates in properties erase_metadata

Delete all annotations in case-file.docx, then erase the author, company and date properties. Apply each step to the file produced by the previous one.

redact_annotations can replace annotation text or, with deleteAll, remove annotations entirely. erase_metadata accepts a comma-separated field list such as Author,Company,CreatedTime, and defaults to All. The chaining instruction matters: each call writes a new file, so an agent that feeds the original into step two throws away step one.

Expected result: a final file produced by the last step in the chain, with every intermediate file name visible in the transcript.

Way 3: Redact before you prompt, and keep the values out of the conversation

Pre-LLM redaction means the sensitive strings never reach a hosted model in the first place. The documents stay on your machine, but the conversation does not. With a cloud-hosted model, what the agent says travels to the model provider, and in redaction work that can include the strings you are removing. “I redacted jane.smith@example.com” has just sent the address you were protecting.

Redact the client name using our standard pattern. Do not repeat the matched text in your reply.

The redaction tools return the saved file name, not a match count. A number in the agent’s reply is its own reading, for example from searching the text before redacting, so treat it as a cross-check and not as tool output. For the strongest setting, run a local model so nothing leaves your perimeter at all, which also gives you document redaction automation that stays local end to end.

Expected result: a saved file name in the reply, with no sensitive strings in the transcript.

Example session (abridged)

You:    Redact every email address in case-file.pdf and tell me which file you saved.
Agent:  I will use this pattern: [\w.+-]+@[\w-]+\.[\w.]+
        -> redact_text(file: case-file.pdf, pattern: ..., replacement: "[REDACTED]")
Agent:  Saved case-file_redacted.pdf.
You:    For case-file.docx, delete all annotations and clear the document properties.
Agent:  -> redact_annotations(file: case-file.docx, deleteAll: true)
        Saved the output file.
        -> erase_metadata(file: the file saved by the previous step, fields: "All")
        Saved the output file.

What are the limits of the evaluation mode?

State this plainly before any run that matters. Without a license the server runs in evaluation mode: one document per process, one redaction, capped at 4 replacements, plus trial badges. A pattern that occurs ten times is redacted four times and left in place six, and the output can look clean. The cap is for demos only. Never ship the output of an unlicensed run, and ask get_license_status first: “What is the license status of the redaction server?”

Whether a result suits a regulation such as GDPR is a decision that stays with you. The server removes what you tell it to remove, in the file you give it.

  • PDF on Linux. On Linux, including the Docker image, redact_image_area and erase_metadata currently fail on PDF files (the PDF engine’s image handling depends on System.Drawing, which .NET supports only on Windows); both work on Word documents, and redact_text works on PDF. Run PDF area redaction and metadata erasure on Windows with dnx.

FAQ

Can Claude redact PII from a PDF without uploading it? Yes. The MCP server runs as a child process of your AI client over local stdio and reads and writes files in folders you configure. Be aware that anything the agent says in chat still goes to a cloud model provider.

Does the agent decide what is sensitive? No. The agent proposes patterns and you review them. The engine applies exactly what it is given.

How do I install it? With the .NET 10 SDK, run dnx GroupDocs.Redaction.Mcp --yes and set GROUPDOCS_MCP_STORAGE_PATH to the folder that holds your documents. Docker is the other option: ghcr.io/groupdocs-redaction/redaction-net-mcp.

Go deeper