Should a law firm test Sonnet 5.5 on document preparation before changing its workflow?

Plan a scoped Sonnet 5.5 document pilot with test material, citation checks, measured review effort and confidentiality decisions before changing workflows.

AI Automation
6 October 2026Updated 06 Oct 20267 min readBukhosi Moyo

Quick Answer

A scoped pilot is appropriate before changing a law firm's document workflow. Compare Sonnet 5.5 and the current process on the same approved samples, checking source fidelity, citations, uncertainty and professional review effort. Begin with fictitious material or samples whose de-identification and processing have been approved. Anthropic's September announcement is a reason to evaluate the model, not evidence of legal accuracy or confidentiality in your selected arrangement.

Key Takeaways

  • Compare the same approved samples and task definitions.
  • Check every citation against the actual source.
  • Record review effort and failures rather than assuming savings.
  • Approve confidentiality and processing before using real matter material.

Want the full breakdown? Scroll below.

People reviewing work together at a desk with laptops
On this pageJump to a section
  1. 1Use the announcement to define a narrow question
  2. 2Approve the material before evaluating the output
  3. 3Scoped legal-document pilot
  4. 4Verify citations rather than rewarding their appearance
  5. 5Review confidentiality for each selected arrangement
  6. 6Work through normal, missing and duplicate samples
  7. 7FAQ about testing Sonnet 5.5 in a law firm
  8. 8Sources

Share this article

Bukhosi Moyo

Growth Partner

Need help growing your company?

We build SEO-first websites and growth systems for South African businesses.

Get Started

A new model announcement can justify an evaluation without justifying a workflow replacement. A law firm needs to know whether document preparation stays faithful to sources, whether the reviewer can spot errors and whether the processing arrangement is acceptable for the material. Vendor benchmarks do not answer those questions for a particular firm.

The proposed pilot below covers a short source summary and a draft preparation note. It keeps legal interpretation and release with the responsible professional. The examples are fictional, and the protocol includes a separate gate for any anonymised real-source sample; no actual pilot results are claimed.

Use the announcement to define a narrow question

Anthropic's 28 September 2026 Sonnet 5.5 announcement describes strengths in well-scoped tasks and polished documents, slides and spreadsheets. The primary page was checked on 6 October. Those are vendor capability claims, not proof of accuracy on legal documents or a measured benefit for your firm. Source: Anthropic Sonnet 5.5 announcement

Ask a narrow evaluation question: can the proposed model prepare a source-faithful draft that a professional can review efficiently under the approved handling arrangement? Do not ask it to decide whether the firm may act, interpret a disputed deadline or provide an unreviewed client opinion.

Check the actual model identifier, interface, account access and available controls before testing. The announcement's availability language does not establish your firm's contract, deployment configuration or professional acceptance. Record those facts as part of the pilot rather than assuming every platform has identical behaviour.

Approve the material before evaluating the output

Begin with entirely invented documents that reproduce the structure of the intended task. Do not copy a real matter and change only its names. Rare events, dates, locations, metadata and linked files may still identify the people or matter.

If the firm wants an anonymised real-source example, the responsible professional and information owner should approve the de-identification, residual disclosure risk, processing destination and access before the sample enters the pilot. A label saying anonymous is not evidence that those decisions were made.

Keep the exact approved sample set versioned. Both the candidate model and baseline receive the same permitted material for the same task. A model that sees additional source context cannot be compared fairly with one that does not, and a reviewer needs to know what was actually supplied.

Scoped legal-document pilot

Use this complete proposed protocol. The initial sample pack is fictional; any approved anonymised real-source sample is recorded as a separate later test condition.

Task A: prepare a factual source summary with citations and explicit uncertainties. Task B: prepare an internal document-preparation note containing only the approved source facts and requested format. Neither task provides legal advice, sends communication or changes a matter record.

Sample Required output Acceptance check Failure or review question
PILOT-A clear fictional source Short summary preserving the stated facts Every factual statement links to supporting source text Unsupported additions fail the source-fidelity check
PILOT-B missing referenced schedule Summary with the missing source clearly identified No invented schedule content or conclusion depending on it A polished complete-looking answer is a failure
PILOT-C contradictory date statements Both attributed statements and an unresolved conflict No automatic selection of the more plausible account Reviewer checks whether uncertainty remains visible
PILOT-D internal and client-permitted fields Internal preparation note using the approved audience rule Excluded fields do not enter the permitted draft Paraphrased internal strategy counts as leakage
PILOT-E duplicated document copies One source-derived fact with duplicate provenance No inflated event count or independent-corroboration claim Similar wording is not enough to establish source identity
PILOT-F unreadable passage or unusable response Explicit hold or partial-coverage statement No guessed text and no false claim of full review Reviewer records the source or interface limitation

For each model or baseline run, store sample version, task wording, model and interface identifier, configuration, input scope, exact output, citation results, unsupported statements, omission findings, confidentiality checks, reviewer corrections and active review time. Record measured effort directly; do not estimate savings from vendor speed or cost claims.

Professional review: inspect every citation and important statement against the source. Identify whether an error is obvious or could survive ordinary review. Record why the draft is accepted, repaired or rejected, including omissions that a fluent answer can hide.

Decision record: compare observed results under the same conditions, identify unresolved risks and state which narrow task, if any, is suitable for a further authorised pilot. Do not infer a firm-wide replacement decision or legal accuracy from a small synthetic sample. Retain the baseline and a clear operational fallback.

The pilot specification is ready when material, task scope, handling arrangement and acceptance rules have approved owners. Completing the specification does not mean the pilot has run or that any model passed it.

Verify citations rather than rewarding their appearance

A citation can look precise while pointing to a passage that does not support the statement. Review the actual relationship: source identity, location, wording and the assertion made. A source mentioning a date does not necessarily support describing it as a deadline.

Check omissions as well as additions. A summary may accurately cite several sentences while leaving out a condition that changes their meaning. The professional should inspect the approved source coverage and task requirements, rather than judge only the apparent polish of the output.

For an existing OpenAI-based baseline, its Structured Outputs guide documents constrained response formats and refusals. This can support the baseline's field checks, but is not evidence that Sonnet uses the same feature or that schema-valid output is legally accurate. Use the candidate interface's actual documented capabilities and review its content independently. Source: OpenAI structured outputs

Review confidentiality for each selected arrangement

If the baseline uses OpenAI, its data-controls guide distinguishes abuse-monitoring logs and application state with endpoint-specific conditions. Those are OpenAI-specific controls; they do not describe Anthropic's selected deployment or establish the firm's confidentiality requirements. Obtain current evidence for each actual provider and interface before using real material. Source: OpenAI data controls

Map sample storage, provider requests, output copies, reviewer workspaces and logs. Have the responsible owners assess the required access, retention and disclosure arrangements. A vendor's general availability or retention capability is not proof that the selected account is configured appropriately.

Keep test results inside the approved audience. Even a de-identified sample can have residual sensitive context, and reviewer notes may describe the original source. Do not publish the outputs or use them in marketing as evidence of legal performance without the relevant authority and factual basis.

Work through normal, missing and duplicate samples

In a hypothetical normal sample, PILOT-A contains a clear factual sequence. The model prepares the requested summary, and the reviewer checks each source link and records actual correction effort. No benefit is claimed until observed evidence exists.

In a missing-source sample, PILOT-B refers to an absent schedule. A useful output identifies the limitation and holds dependent statements. Inventing plausible schedule content fails the pilot even if the draft reads smoothly.

In a duplicate sample, PILOT-E contains two copies of the same verified document. The output should avoid treating repeated wording as two independent facts or sources. If the versions conflict, the reviewer needs that difference preserved rather than resolved automatically.

FAQ about testing Sonnet 5.5 in a law firm

Does a vendor benchmark establish legal-document accuracy?

No. Evaluate the exact task, source fidelity, reviewability and handling arrangement. Benchmarks can inform the reason to test but cannot replace evidence from the firm's authorised evaluation.

Can we assume an anonymised matter sample is safe to upload?

No. Have the responsible owners approve the de-identification and selected processing arrangement. Start with independently invented fixtures when those questions are unresolved.

What should justify changing the current workflow?

A documented decision based on observed comparable results, remaining risks, professional review and operational controls. A small successful test can support a bounded next step, not an unsupported claim that every matter task is suitable.

If your business needs help defining this process, explore AI agents for law firms, the wider AI automation services, and our custom-agent workflow guide. The agents and automation comparison and custom AI agent glossary explain the terms. To discuss your records and approval rules, get in touch.

Sources

Share this article

Bukhosi Moyo

Written by

Bukhosi Moyo

CEO & Founder

Bukhosi is the founder and lead SEO strategist at Symaxx. He architects search-first digital systems for South African businesses, combining technical engineering with commercial strategy to build long-term organic assets.

Feedback

Was this helpful?

Tell us how this article felt in one click.

Back to Insights

Need help executing this strategy?

Our team turns these insights into revenue-generating search architectures for your business.