Introduction

A colleague sends two revisions of a contract and asks you to diff them. You drop both into your comparison service, the result comes back, and everything looks normal. What you did not see is that one of those documents carried a linked image pointing at a URL, and your server contacted that host the moment the file was opened. Nothing about the output tells you it happened.

This is not a defect - it is what loading a document faithfully means. An OOXML file can reference an image that lives on a web server rather than inside the package, and both Word and any library that loads the document properly resolve that reference. GroupDocs.Comparison for .NET exposes two properties on LoadOptions that let you decide whether it does: SkipExternalResources and WhitelistedResources.

Between them they give three configurations, and this article compares all three - the permissive default, blocking everything, and blocking everything except named references. By the end you will know which to pick for a given document source, and the two mistakes that make these settings look as though they do not work.

💡 Full working example: block-external-resources-on-document-load-dotnet - a runnable console project that serves the referenced images itself and logs every request, so you can watch each setting take effect.

Where External References Hide

Before choosing a setting, it is worth knowing what you are choosing about. A .docx carries external references in two distinct places, and they are easy to miss because neither is visible in the document text.

The first is a relationship in word/_rels/document.xml.rels carrying TargetMode="External" and an absolute URL. The image shows up in the body as a drawing that points at the relationship by ID, so the URL itself never appears near the content it affects.

The second is an INCLUDEPICTURE field code in the document body, holding its URL inside a field instruction. Word resolves it when the page renders; a comparison library resolves it when the document loads.

Both mechanisms respect the two load options discussed below, which matters because a document can use either or both. A reference you spotted in the relationships file is not proof there is no second one in a field code.

Approach 1: The Default - References Resolved

SkipExternalResources defaults to false, so a document loaded without configuration has its remote references resolved:

LoadOptions loadOptions = new LoadOptions
{
    SkipExternalResources = false
};

using (Comparer comparer = new Comparer(sourcePath, loadOptions))
{
    comparer.Add(targetPath, loadOptions);
    comparer.Compare(outputPath);
}

This gives the highest fidelity: the compared documents contain everything they reference, exactly as Word would render them. For documents your own application or templates produced, where every reference URL points at infrastructure you operate, it is the right choice - and a missing linked image could make the comparison actively misleading.

The cost is that every reference is contacted, whoever put it there. There is also a timing cost that has nothing to do with trust: a reference URL that no longer resolves makes loading wait out the full connection attempt, on every single comparison.

Approach 2: Block Every External Resource

One property turns remote reference resolution off for that document:

LoadOptions loadOptions = new LoadOptions
{
    SkipExternalResources = true
};

using (Comparer comparer = new Comparer(sourcePath, loadOptions))
{
    comparer.Add(targetPath, loadOptions);
    comparer.Compare(outputPath);
}

No request is issued. The referenced images are absent from the result, and - this is the part worth being clear about - nothing else changes. The setting governs what gets loaded, not how differences are found, so textual and structural changes between the two documents are detected exactly as before. The one thing you lose is the ability to detect a change inside a referenced image, which was never loaded.

This is the configuration to treat as your baseline for documents you did not create: user uploads in a web application, files received by email, anything compared on a build agent where an outbound request is rarely intended. It is all or nothing, though - a linked image you actually wanted is blocked along with the rest, and the result simply lacks it without announcing the fact.

Approach 3: Block Everything Except Named References

The third configuration is the one that rewards a close reading. WhitelistedResources takes a List<string> and is consulted only when SkipExternalResources is true:

LoadOptions loadOptions = new LoadOptions
{
    SkipExternalResources = true,
    WhitelistedResources = new List<string> { "includepicture-field.png" }
};

using (Comparer comparer = new Comparer(sourcePath, loadOptions))
{
    comparer.Add(targetPath, loadOptions);
    comparer.Compare(outputPath);
}

The entries are URL fragments, not file names. Each is matched against the reference URL, and a match anywhere in it admits that resource. That is what makes the whitelist portable: "includepicture-field.png" admits the image whatever scheme, host and path precede it, so the same list works in development and production without rewriting.

The same property cuts the other way. A short or generic fragment - logo.png, or worse, .png - can match references you never intended to allow. Pick a fragment specific enough to identify the one resource you meant.

In the reference sample, this configuration fetches the whitelisted image and leaves a second referenced image, which no entry covers, blocked. The request log shows three requests where the permissive default produced five, and names only the whitelisted file.

Which Configuration Should You Use?

Match the setting to where the document came from. Documents your own application or templates generated can keep the default, because every reference URL points at infrastructure you already operate. Anything arriving from outside - user uploads, email attachments, third-party files - warrants SkipExternalResources = true. Add a narrow WhitelistedResources fragment only when one trusted reference genuinely has to resolve.

Comparing the Three

Concern Default Skip all Skip + whitelist
Properties to set 0 1 2
Outbound requests all references none whitelisted only
Per-reference control no no yes
Dead URL costs load time yes no whitelisted only
Best for documents you produced documents from anywhere else trusted templates among untrusted content

The decision follows document provenance rather than performance. Documents your systems generated can keep the default. Documents from outside warrant blocking. Whitelist at the point where one specific reference genuinely has to resolve - a corporate template pulling its header image from an internal URL, say, among reports whose authors pasted in images from wherever they liked.

The Two Mistakes

Both of these produce the same symptom: you set the option, and it appears to do nothing.

A whitelist without the switch. WhitelistedResources is consulted only when SkipExternalResources is true. Set on its own, it does nothing at all - there is no blocking for it to make an exception to. If a whitelist looks ignored, check this first.

Options on the source only. This is the subtler one. Load options describe how one document is loaded. The Comparer constructor takes the options for the source; every Add() call takes the options for that target:

using (Comparer comparer = new Comparer(sourcePath, loadOptions))
{
    comparer.Add(targetPath, loadOptions);
    comparer.Compare(outputPath);
}

Pass them to the constructor and forget the Add() call, and the source is protected while every target still fetches its references. The comparison succeeds, the result looks plausible, and half your documents are still reaching out to the network. Where source and target need different handling, pass separate LoadOptions instances - that is precisely why the API takes them per document.

Verifying It Actually Worked

A blocked resource leaves almost no trace. The output document is missing an image, which looks much like a document that never had one. Reading the result file is therefore a poor way to confirm the setting took effect.

Watch the serving side instead. The reference sample takes this approach deliberately: it starts a small HTTP listener on a free loopback port, writes its demo documents pointing at that port, and logs every request it receives, printing the count per comparison. Five requests, then zero, then three. A network trace against your real document sources gets you the same confidence.

Conclusion

Three configurations, one decision rule: let documents you generated keep the default, set SkipExternalResources = true for everything else, and whitelist a narrow URL fragment only where a specific trusted reference still needs to resolve.

Then check the two things that silently undo the work - a whitelist without SkipExternalResources = true, and options passed to the Comparer constructor but not to every Add() call - and verify from the serving side rather than the output file.

Additional Resources