When a Law Firm Automation Runs Twice: Preventing Duplicate Tasks, Emails and Records
A law firm automation can sometimes run the same task twice. Learn how duplicate triggers, retries, record checks and reconciliation can help prevent duplicate work.
A law firm automation can sometimes run the same task twice. Learn how duplicate triggers, retries, record checks and reconciliation can help prevent duplicate work.
An AI tool can appear the same after an upgrade. The button will be in the same location, the workflow can be the same, and employees will not see anything unusual when they launch the tool for the first time. However, under the hood, the provider may have switched out the model, altered the processing of the prompt, tweaked the workflow, or made changes to some of the other tools used in conjunction with it.
In terms of a law firm, the real question is not whether the new version is better but whether it still does the job that the firm has already vetted. It is possible that an upgrade which enhances the performance of one task will change the performance of another task.
That is where evaluations, also known as "evals," come in. According to Anthropic, an eval is a test whereby an AI system gets a certain input and its output is then evaluated against certain criteria. The regression tests are done to make sure that the system is still performing well on tasks that it had performed successfully before.
This does not necessarily mean developing an extensive testing process for each change made to an AI system for a law firm. A few examples of previously reviewed work would be a good place to start.
Use work the firm has already approved
Good examples may be right there in the workflow of the firm.
For instance, if an AI system is employed to summarize the notes of meetings, then the firm will have already checked several summaries and will know what a good outcome would look like. That can be used as the example set for that firm. This can be applied to document summarization, intake classification, e-mail classification or draft generation, as long as those are good examples.
This is more useful than depending solely on demonstrations. Although the sample question posed by a vendor might demonstrate that the system can accomplish the task, it does not tell the company whether an update has altered the process of how its work is done. The previously reviewed matter summary, document classification, or draft gives the team something more relevant to compare.
Similarly, the guidance provided by Anthropic stresses the importance of choosing tasks that mimic actual usage and include failure cases or manual checks encountered by the team.
Comparison of the old results and new
The comparison between the old and new will bring to light some changes that might be overlooked by memory. Pick any example of the work that has been approved before and enter it into the new system and then do a comparison with the old one. It may just happen that the only thing different about the two results is the way of presenting them but they have the same significant information.
The new one misses the deadline or treats an important aspect in a different manner.
The comparison is helpful even if the answer is "nothing important happened." The company will have something real to demonstrate instead of one person's opinion that the system looks as though it's working the way it did previously.
It can also assist in distinguishing trivial differences from significant ones. It's no cause for worry if the differences in style and conciseness don't impact on accuracy of the information required by the workflow process.
Look out for regressions that appear not to be failures
Regression failures don't always come with errors.
Think of an AI solution that classifies incoming emails into practice groups. Before the upgrade, it would reliably identify a certain type of request. Now, after the upgrade, it provides classification, thus the process seems to be going smoothly, but the emails of the same type are not being classified as reliably as before.
No crash has occurred. No error message was reported. This difference might be revealed only by examining a number of outcomes.
This is one of the reasons why regression testing is important. Anthropic differentiates regression assessments from capability assessments based on their goals: regression testing is about whether the system continues to do tasks that it used to do successfully.
For the law firm, those existing tasks can get as much priority as the new feature that the provider wants to promote.
Determine what has been changed before determining what is to be tested.
The volume and nature of the testing may differ according to the modification.
The provider who replaces the underlying model is different from the one who makes a typo in the prompt. Modification of the prompt will affect the output even without modification of the model. Modification of the workflow will affect the processes before and after the AI produces its answer. Adding a connection or a data source could be an additional variable.
It would be wise to make a note before the test on what changes were made: model, prompt, workflow, connection/data source, or post-processing?
These notes will provide clues for the reviewer where to look. If the prompt changed, the examples which relied much on the previous instructions should receive special consideration. If a step of the workflow was added/changed, it would not suffice just to look at the AI's answer.
Look beyond the problems you already know about
Just testing the old problems that cause failure can lead to a false sense of security, since an update may cause a completely new problem to arise.
The output may now contain data not covered by the source document. The date may be processed differently, the document assigned an incorrect classification or an essential field removed. In a bigger process flow, the AI-generated output may be correct but further processing may fail.
However, just one strange output is not necessarily proof of failure due to the update, as there can be variation between the outputs generated by AI. The critical point is if the problem is consistent across tests. The section of Anthropic on agent evaluations makes this distinction relevant in cases where a team needs to identify if there is a true regression in the changes made or simply variation.
Provide a few challenging examples
The easy case is critical, but it should not be the only one that is tested by the company.
There are several challenges that can be included in the test set such as a document that is in an unusual format, an incomplete request, names that are similar within a matter, a very long document or an email that is difficult to categorize.
In addition, it might be helpful to have some instances where the right course of action is to seek clarification instead of providing an answer with full confidence. This is not about generating a vast list of weird situations. Just a few representative cases could demonstrate whether the updated system performs differently when the information it gets is not completely clean.
According to Anthropic, clear tasks and success criteria need to be defined so that one could evaluate the success of the AI system. This approach could also be used in a smaller test set for a law firm.
Someone should own the decision after testing
Conducting the tests is only half of the story. The firm must also identify who will be responsible for making the decision whether to move ahead with the updated version.
Take the case of the update that performs excellently in most of the examples presented by the firm but gives a different result in one particular document that is vital for one practice group. There should be someone who decides whether that is acceptable or not.
The person may be an IT lead, department head, knowledge management expert, or someone else as per the way the AI process is being managed within the company. What matters is that the responsibility is defined.
The report that will be generated will not need to be complex at all. It may just state that the testing was done, the comparison of old and new results was done, the regression was analyzed, and a decision about continued usage was made. In case of any problem, the report can prove that usage was stopped until the problem was resolved.
Keep the test set small enough to use
The testing procedure becomes much less useful when it is so extensive that no one wishes to perform it.
For regular updates, it may be sufficient to conduct a smaller subset of tests including only the most frequently used AI-aided tasks, especially those that created problems before. Anthropic states that the team may begin with a fairly small subset of tasks and increase their scope of testing as they get better acquainted with the system.
A company may retain a controlled set that would include the initial input, the expected or approved output, the previous outputs, the known edge cases and the notes on the previous failure. With each update of the AI tool, the team will have something to rely on instead of conducting the testing process from scratch.
Retaining the old results becomes particularly useful. In case the system gets updated several months later, the reviewers may refer to the old update and use the tests that were useful for them.
A practical way to record the results
There is no need to develop a specific testing platform for every internal workflow. A plain record will contain sufficient information to simplify further reviews.
Every test must include ID, original input, original approved output, current output, changes, reviewer, outcome, and date of review. In such a way, the reviewer will have certain information that may help to understand whether the issue is something new or just a case of something that the company experienced earlier.
The same information is useful if a provider publishes another update. Instead of asking what had been tested last time, the team is able to refer to its previous experience and choose the appropriate tests.
Some questions to be answered before publishing the update
Prior to returning an updated AI workflow into use, the reviewer must be able to answer to the following questions:
What was changed? Model? Prompt? Workflow itself? Tools? Something else? Which approved samples were tested again? The cases must be based on real-life scenarios involving work which has been analyzed by the company previously.
What was unique about the new result when compared to the previous result? Consider only those things that could impact the firm's activities, but don't make every little formatting difference into an issue.
Were there any new mistakes made? Consider both typical and more complex mistakes. Was the workflow still possible from start to finish? If the AI is used within some bigger process, look at the things that happened before and after the AI's output. Who made sure the results were fine? Someone must take responsibility for the decision about whether the updated version can be accepted. What was the decision? Keep a record of this so you have something to refer to when the next update comes around.
An update is an occasion to inspect, not to worry
AI tools will evolve. It is not necessary for a firm to consider each update as an excuse to cease applying a particular workflow, but it must not think that everything is the same if an interface remains unchanged.
A limited number of authorized examples, comparisons between old and new results and some tough examples will help detect changes long before they will become a usual thing. Having a reviewer and keeping results will make the final decision more obvious and will give a firm an opportunity to have a starting point for the next update.
The process of testing can be made more valuable in the course of time. Something that is not working now can become a test case tomorrow, and something that is authorized can serve as an example of comparison during future changes.
AKAVEIL TECHNOLOGIES works with law firms on AI implementation and testing, including reviewing AI workflows, establishing appropriate controls and testing changes before they become part of everyday operations.
To discuss AI implementation and testing for your law firm, visit akaveil.com or write to us at info@akaveil.com.
About Ariel Perez
Ariel Perez is the Founder and CEO of AKAVEIL Technologies, where he works with law firms on Microsoft 365, cybersecurity, cloud infrastructure and AI readiness. His work often involves reviewing the systems, permissions and access controls that sit underneath tools such as Microsoft Copilot and other AI workflows.
With nearly two decades of experience in IT, cloud infrastructure and cybersecurity, Ariel helps law firms understand how new technology interacts with the environment they already have in place, including SharePoint, OneDrive, Teams, Microsoft Entra and other parts of Microsoft 365.
For questions about Microsoft 365, AI readiness or cybersecurity for your law firm, write to us at info@akaveil.com.
About the Author
Ariel Pérez
Founder & CEO of AKAVEIL Technologies, Ariel brings nearly two decades of expertise in IT, cloud infrastructure, and cybersecurity exclusively for law firms. He specializes in Microsoft 365, Azure Virtual Desktop, and AI-driven automation, helping legal organizations transition from legacy systems to modern cloud platforms. Ariel's deep understanding of legal workflows and hands-on technical approach makes him a trusted advisor for law firm leadership seeking to enhance security, compliance, and operational efficiency.
Automation & AI Pipelines
Workflow automation and AI document processing for law firms using n8n and Microsoft automation services.
Ready to Secure Your Law Firm?
Let AKAVEIL help you implement comprehensive cybersecurity solutions.
Continue Reading
Explore more insights on legal technology and IT solutions.