The Problem
You upload a 300-page facility agreement. Ask the AI to review it. The AI confidently produces analysis—but it's subtly wrong. Provisions in the later sections are misrepresented or missing entirely.
You didn't hit an error message. The AI didn't tell you it couldn't handle the document. It just... degraded quietly.
This is context window failure. It is hard to catch because nothing tells you it has happened.
This page is about the length of a document. It does not make a client's document safe to upload. What you may put into an AI tool depends on the plan your firm has, its agreement with the vendor and your firm's policy: Claude Projects for Client Matters sets out the checks for one tool. To test the patterns below, use a precedent or a published agreement that contains no client information.
Understanding Context Windows
Every AI model has a context window: a limit on how much text it can take into account at once, including its own answer. Anthropic's documentation calls it the model's "working memory".
The figure depends on the model, and it changes. Vendors replace their models often and the context window changes with them, so this page quotes no figure for any named model. Find the current one in the vendor's documentation for the model your firm uses. For Claude, for example, Anthropic publishes it in How large is the context window on paid Claude plans? and in its developer documentation on context windows.
But here's the critical insight: a document that fits is not a document that is read well. Anthropic's context windows page says that "more context isn't automatically better" and that "As token count grows, accuracy and recall degrade" (checked 29 September 2026).
The 60% Rule
Never use more than 60% of the stated context window for production legal work.
This is a working margin, not a measured threshold. No test results have been published for it. Models tend to get less reliable as the context window fills, so start at 60% and find the right figure for your model by testing it on documents where you already know the answers.
Why leave a margin?
- Quality degradation: Anthropic says of its own models that accuracy and recall degrade as the token count grows (quoted above)
- Room for the answer: The context window has to hold the AI's response and any system prompt as well as your document
Neither reason fixes the figure at 60%. It is a suggested place to start.
In practice:
Take the context window the vendor states for your model and multiply it by 0.6.
- Stated window of 200,000 tokens: use no more than 120,000
- Stated window of 1,000,000 tokens: use no more than 600,000
These two figures are worked examples of the arithmetic, not the limits of any particular model.
What Can Happen When You Exceed the Limit
Subtle failures (hardest to catch):
- Provisions in later sections are ignored or misrepresented
- Cross-references aren't followed completely
- Comparative analysis weighs earlier sections more heavily
- Summaries omit details from the end of documents
Obvious failures (easier to catch):
- Explicit token limit errors
- Truncated responses
- Generic/vague outputs
- "Unable to process" messages
The subtle failures are the dangerous ones for legal work, because nothing tells you they have happened.
Context Architecture Patterns
Pattern 1: Chunking with Context Preservation
Instead of sending the entire document, break it into meaningful chunks that preserve context.
Bad chunking (arbitrary splits):
Chunk 1: Pages 1-100
Chunk 2: Pages 101-200
Chunk 3: Pages 201-300
Good chunking (logical sections):
Chunk 1: Clause 1 (Definitions and interpretation) + Clause 2 (Warranties)
Chunk 2: Clause 3 (Undertakings) + Clause 4 (Conditions precedent)
Chunk 3: Clause 5 (Indemnities) + Clause 6 (General) + Schedules
Preserve context by:
- Including relevant definitions in each chunk
- Summarising previous chunks when analysing later ones
- Maintaining cross-references within chunks
Pattern 2: Progressive Disclosure
Start broad, then zoom in on areas of interest.
Pass 1 - Structural Analysis (Whole document, light analysis):
Review this 300-page agreement and provide:
- Table of contents
- Key provisions by category
- Notable unusual provisions
- Sections requiring detailed review
Pass 2 - Targeted Deep Dive (Selected sections, detailed analysis):
Now review Clause 8 (Indemnities) in detail:
[Include only Clause 8 + relevant definitions]
- Scope of the indemnities
- Carve-outs and limitations
- Procedures and timelines
- Comparison to market standard
This two-pass approach stays well within context limits while still achieving comprehensive coverage.
Pattern 3: Hierarchical Summarisation
For truly massive document sets, build summaries hierarchically.
Level 1: Individual document summaries (each under 60% limit)
Level 2: Combine summaries for comparative analysis
Level 3: Full synthesis with references to source documents
Example - Due Diligence:
Level 1: Summarise each of 50 contracts individually
Level 2: Group summaries by category (employment, supplier, IP)
Level 3: Cross-contract analysis of key terms
Level 4: Risk assessment and recommendations
Each level stays within context limits. Final analysis can reference source documents by name for spot-checking.
Real-World Application: M&A Due Diligence
Scenario: Reviewing 200 contracts in a data room
Naive approach (fails): Upload all 200 contracts, ask for analysis → Context window exceeded or severe quality degradation
Context-aware approach:
Step 1 - Categorisation (one-by-one):
For each contract:
- Document type
- Parties
- Key terms (dates, amounts, obligations)
- Material provisions
Step 2 - Category analysis (batched):
Group by category:
- All employment agreements → analyse together
- All supplier contracts → analyse together
- All IP licences → analyse together
Step 3 - Cross-cutting analysis (summary-based):
Using summaries from Steps 1-2:
- Change of control provisions across all contracts
- Termination rights patterns
- Liability caps and limitations
- Consent requirements
Step 4 - Deep dives (targeted):
For flagged contracts, perform detailed review
with full text + relevant comparisons
This approach keeps each individual AI call well within limits while achieving comprehensive coverage.
Token Counting
Before sending a document to an AI, estimate its token count.
Rules of thumb (checked 29 September 2026):
Each vendor counts tokens in its own way, and the ratio can change when a model changes.
- Google: "For Gemini models, a token is equivalent to about 4 characters. 100 tokens is equal to about 60-80 English words." (Understand and count tokens)
- Anthropic: "For Claude, a token approximately represents 3.5 English characters, though the exact number can vary depending on the language used." (Glossary)
- Anthropic also says that its newer models produce roughly 30 per cent more tokens from the same text than its earlier ones. (Token counting)
When in doubt, use the estimate that gives the higher token count.
For a precise count:
Use the counter the vendor provides for the model you use. Anthropic documents one for developers on its Token counting page. If nobody at your firm can run it, rely on the rule of thumb and the 60% margin.
Quick check:
300-page agreement at about 500 words a page
≈ 150,000 words
≈ 190,000 to 250,000 tokens (at 60 to 80 words per 100 tokens)
Count the words in your own document rather than assuming 500 a page, because layout varies.
With a stated window of 200,000 tokens, the working margin is 120,000. This document is over the margin on either estimate, so it needs chunking or progressive disclosure. With a stated window of 1,000,000 tokens, the margin is 600,000 and the document is inside it.
Warning Signs You've Exceeded Capacity
- Vague responses: AI becomes less specific about later sections
- Repeated phrases: Copy-paste patterns suggest shallow processing
- Missed details: Provisions you know exist aren't mentioned
- Inconsistent depth: Earlier sections get detailed analysis, later sections get summaries
- Generic language: "May contain", "typically includes" instead of specific findings
If you see these patterns: Your context may be too large. Chunk it and compare the results.
Advanced: Context Window Budgeting
For complex prompts, budget your context window:
Total budget (60% of a stated window of 200,000 tokens): 120,000 tokens
Allocations:
- System prompt & instructions: 2,000 tokens
- Examples & formatting guidance: 3,000 tokens
- Document content: 110,000 tokens
- Response buffer: 5,000 tokens
The figures are an illustration. Set your own from the window your vendor states.
Budgeting this way means a limit error does not take you by surprise.
Special Case: Long-Form Document Generation
When the AI needs to generate a long document (not just analyse), context fills up fast.
Problem: You want a 50-page contract. The AI starts strong but degrades in later sections.
Solution - Iterative Generation:
- Outline first: Generate full structure/table of contents
- Clause by clause: Generate each clause separately
- Cross-reference pass: Ensure consistency across clauses
- Final assembly: Combine with human review
Each generation step stays within context limits.
How This Fits with the Rest of the Methodology
Context architecture affects everything:
- Testing Checklist: Add "Context window budget verified" as a check item
- Tracer Bullets: Test with documents at 60% threshold before going larger
- Validation Layers: Automated token counting before human review
Practical Implementation Checklist
Before sending any document to AI:
- Estimate token count (conservative)
- Check against 60% threshold for your model
- If exceeds threshold, choose chunking strategy
- Document which chunking approach you used
- Verify outputs don't show degradation warning signs
- Spot-check analysis of later sections
Common Mistakes
Mistake 1: "The vendor states a limit, so I can use all of it"
A stated limit tells you what the model will accept, not how well it will read it. Leave a margin for professional work.
Mistake 2: "I didn't get an error, so it must have worked"
Context degradation is silent. The AI won't tell you it's struggling.
Mistake 3: Arbitrary chunking
Don't break documents mid-section. Chunk at logical boundaries (clauses, parts, schedules).
Future-Proofing
Context windows change when vendors release new models. Whether the same margin will suit a later model is not known until it is tested.
When the model behind your tool changes, test again. Run documents where you know the answers at several sizes and see where the quality falls. You may find that the right margin for that model is 50% or 70%. Set it from your own results.
Getting Started
- Estimate the token count of a typical document you work with
- Compare to 60% threshold for your preferred model
- If you're over the limit, design a chunking or progressive disclosure strategy
- Test it on a sample document where you know the right answers
- Document your approach so others can replicate it
Context architecture isn't glamorous, but it's the difference between AI outputs you can trust and AI outputs that quietly fail in critical ways.
Need help designing context-aware workflows for complex document analysis? The consulting page describes what is offered.