A new model announcement can justify an evaluation without justifying a workflow replacement. A law firm needs to know whether document preparation stays faithful to sources, whether the reviewer can spot errors and whether the processing arrangement is acceptable for the material. Vendor benchmarks do not answer those questions for a particular firm.
The proposed pilot below covers a short source summary and a draft preparation note. It keeps legal interpretation and release with the responsible professional. The examples are fictional, and the protocol includes a separate gate for any anonymised real-source sample; no actual pilot results are claimed.
Use the announcement to define a narrow question
Anthropic's 28 September 2026 Sonnet 5.5 announcement describes strengths in well-scoped tasks and polished documents, slides and spreadsheets. The primary page was checked on 6 October. Those are vendor capability claims, not proof of accuracy on legal documents or a measured benefit for your firm. Source: Anthropic Sonnet 5.5 announcement
Ask a narrow evaluation question: can the proposed model prepare a source-faithful draft that a professional can review efficiently under the approved handling arrangement? Do not ask it to decide whether the firm may act, interpret a disputed deadline or provide an unreviewed client opinion.
Check the actual model identifier, interface, account access and available controls before testing. The announcement's availability language does not establish your firm's contract, deployment configuration or professional acceptance. Record those facts as part of the pilot rather than assuming every platform has identical behaviour.
Approve the material before evaluating the output
Begin with entirely invented documents that reproduce the structure of the intended task. Do not copy a real matter and change only its names. Rare events, dates, locations, metadata and linked files may still identify the people or matter.
If the firm wants an anonymised real-source example, the responsible professional and information owner should approve the de-identification, residual disclosure risk, processing destination and access before the sample enters the pilot. A label saying anonymous is not evidence that those decisions were made.
Keep the exact approved sample set versioned. Both the candidate model and baseline receive the same permitted material for the same task. A model that sees additional source context cannot be compared fairly with one that does not, and a reviewer needs to know what was actually supplied.
Scoped legal-document pilot
Use this complete proposed protocol. The initial sample pack is fictional; any approved anonymised real-source sample is recorded as a separate later test condition.
Task A: prepare a factual source summary with citations and explicit uncertainties. Task B: prepare an internal document-preparation note containing only the approved source facts and requested format. Neither task provides legal advice, sends communication or changes a matter record.
| Sample | Required output | Acceptance check | Failure or review question |
|---|---|---|---|
| PILOT-A clear fictional source | Short summary preserving the stated facts | Every factual statement links to supporting source text | Unsupported additions fail the source-fidelity check |
| PILOT-B missing referenced schedule | Summary with the missing source clearly identified | No invented schedule content or conclusion depending on it | A polished complete-looking answer is a failure |
| PILOT-C contradictory date statements | Both attributed statements and an unresolved conflict | No automatic selection of the more plausible account | Reviewer checks whether uncertainty remains visible |
| PILOT-D internal and client-permitted fields | Internal preparation note using the approved audience rule | Excluded fields do not enter the permitted draft | Paraphrased internal strategy counts as leakage |
| PILOT-E duplicated document copies | One source-derived fact with duplicate provenance | No inflated event count or independent-corroboration claim | Similar wording is not enough to establish source identity |
| PILOT-F unreadable passage or unusable response | Explicit hold or partial-coverage statement | No guessed text and no false claim of full review | Reviewer records the source or interface limitation |
For each model or baseline run, store sample version, task wording, model and interface identifier, configuration, input scope, exact output, citation results, unsupported statements, omission findings, confidentiality checks, reviewer corrections and active review time. Record measured effort directly; do not estimate savings from vendor speed or cost claims.
Professional review: inspect every citation and important statement against the source. Identify whether an error is obvious or could survive ordinary review. Record why the draft is accepted, repaired or rejected, including omissions that a fluent answer can hide.
Decision record: compare observed results under the same conditions, identify unresolved risks and state which narrow task, if any, is suitable for a further authorised pilot. Do not infer a firm-wide replacement decision or legal accuracy from a small synthetic sample. Retain the baseline and a clear operational fallback.
The pilot specification is ready when material, task scope, handling arrangement and acceptance rules have approved owners. Completing the specification does not mean the pilot has run or that any model passed it.
Verify citations rather than rewarding their appearance
A citation can look precise while pointing to a passage that does not support the statement. Review the actual relationship: source identity, location, wording and the assertion made. A source mentioning a date does not necessarily support describing it as a deadline.
Check omissions as well as additions. A summary may accurately cite several sentences while leaving out a condition that changes their meaning. The professional should inspect the approved source coverage and task requirements, rather than judge only the apparent polish of the output.
For an existing OpenAI-based baseline, its Structured Outputs guide documents constrained response formats and refusals. This can support the baseline's field checks, but is not evidence that Sonnet uses the same feature or that schema-valid output is legally accurate. Use the candidate interface's actual documented capabilities and review its content independently. Source: OpenAI structured outputs
Review confidentiality for each selected arrangement
If the baseline uses OpenAI, its data-controls guide distinguishes abuse-monitoring logs and application state with endpoint-specific conditions. Those are OpenAI-specific controls; they do not describe Anthropic's selected deployment or establish the firm's confidentiality requirements. Obtain current evidence for each actual provider and interface before using real material. Source: OpenAI data controls
Map sample storage, provider requests, output copies, reviewer workspaces and logs. Have the responsible owners assess the required access, retention and disclosure arrangements. A vendor's general availability or retention capability is not proof that the selected account is configured appropriately.
Keep test results inside the approved audience. Even a de-identified sample can have residual sensitive context, and reviewer notes may describe the original source. Do not publish the outputs or use them in marketing as evidence of legal performance without the relevant authority and factual basis.
Work through normal, missing and duplicate samples
In a hypothetical normal sample, PILOT-A contains a clear factual sequence. The model prepares the requested summary, and the reviewer checks each source link and records actual correction effort. No benefit is claimed until observed evidence exists.
In a missing-source sample, PILOT-B refers to an absent schedule. A useful output identifies the limitation and holds dependent statements. Inventing plausible schedule content fails the pilot even if the draft reads smoothly.
In a duplicate sample, PILOT-E contains two copies of the same verified document. The output should avoid treating repeated wording as two independent facts or sources. If the versions conflict, the reviewer needs that difference preserved rather than resolved automatically.
FAQ about testing Sonnet 5.5 in a law firm
Does a vendor benchmark establish legal-document accuracy?
No. Evaluate the exact task, source fidelity, reviewability and handling arrangement. Benchmarks can inform the reason to test but cannot replace evidence from the firm's authorised evaluation.
Can we assume an anonymised matter sample is safe to upload?
No. Have the responsible owners approve the de-identification and selected processing arrangement. Start with independently invented fixtures when those questions are unresolved.
What should justify changing the current workflow?
A documented decision based on observed comparable results, remaining risks, professional review and operational controls. A small successful test can support a bounded next step, not an unsupported claim that every matter task is suitable.
If your business needs help defining this process, explore AI agents for law firms, the wider AI automation services, and our custom-agent workflow guide. The agents and automation comparison and custom AI agent glossary explain the terms. To discuss your records and approval rules, get in touch.

