Adding genAI-assisted techniques to SOLVE-IT
Weaknesses and mitigations for querying case data with generative AI
Introduction
Digital forensic practitioners have started using genAI-assisted tools to perform analysis tasks and event reconstruction on a set of extracted artefacts. Done correctly, this can help find the most relevant evidence and activities quickly in hundreds of thousands of data points. However, using AI that is not sufficiently guarded and grounded to answer investigative questions (e.g., who was where when) can produce results that are incomplete, inaccurate, or misinterpreted.
This article shows how SOLVE-IT helps address weaknesses specifically in AI-driven evidence analysis and event reconstruction, to help digital forensics practitioners and tool developers mitigate them.
This article covers three topics:
- The scope of “AI-assisted” within an investigative workflow
- Defining the technique: “Use an AI-based prompt for interrogating case data”
- Weakness and mitigation enumeration for the technique
Scope of “AI-assisted”
Section 5.5 of the original paper introducing SOLVE-IT (Hargreaves, van Beek and Casey, 2025)1 describes one of the applications for SOLVE-IT as the structured consideration of the potential for AI use in digital forensics. This is also now an exemplar for the use of the SOLVE-IT-X extensions framework to build additional custom content on top of the core knowledge base. This covers examples such as detecting if an image contains synthetic content (DFT-1176), or extracting artefacts from chat apps (DFT-1072).
However, what this lacks is a technique to support one of the main emerging examples of using AI in digital forensics, which is the use of a generative AI powered interface to allow an investigator to query the case data with natural language. We have now captured this in SOLVE-IT as the technique “Use an AI-based prompt for interrogating case data”.
DFT-1203: Use an AI-based prompt for interrogating case data
DFT-1203 is defined as “Using an AI-based chat interface to query the data within a case”, and the details also describe that “Behind this interface this could include querying artefacts extracted into a database and/or running tools dynamically.”
This technique was developed using the TRWM workflow, which forces the enumeration of the outputs (or results) of the technique being developed, then considers in a structured way the weaknesses associated with those results, followed by mitigations. In this example, so far, we have enumerated four types of results, which are shown in Table 1.
| Result | Current ontological mapping | Example queries that could be issued |
|---|---|---|
| An assertion that an event happened | HypothesisedEvent | Did John Smith go to this address last Thursday? Was there a data breach from this device? |
| The subset of artefacts deemed relevant | ArtifactSet | Can you find me any artefacts that relate to WiFi hacking? |
| An assertion of a link between entities | HypothesisedRelationship | Was this computer connected to that USB stick? Do Alice and Bob know each other? |
| An explanation of the results presented | uco-core:Annotation | [not a query, but the attached explanation of “reasoning”] |
Table 1: Digital Forensic Technique Results (DFTRs) that were used to develop the technique, along with examples of queries that may produce that type of result.
How this technique fits in with current tools and practices is illustrated in a figure available in the interactive SOLVE-IT Workflow builder here (direct link to the rendered diagram), with a simplified version shown below in Figure 1, and a small but visible under zoom version in Figure 2.
This highlights that:
a) the results from using AI in this way can only be as effective as the tooling supporting it, and
b) this AI-driven technique becomes a critical point of failure, with potentially all queries and results moving through this process.
Therefore, understanding the potential weaknesses of this technique and identifying mitigations is essential.


Weakness enumeration
The weaknesses and mitigations here were developed using the TRWM workflow (documented here), which is also available as a helper app at https://trwm.hargs.co.uk/.
- T – Technique description (covered earlier)
- R – Results (covered in the previous section)
- W – Weaknesses (an enumeration using a systematic process considering how the results identified above could be wrong, i.e. what is the failure effect)
- M – Mitigations (what could be put in place to reduce the risk of specific weaknesses)
There are also specific variations of the TRWM workflow: TRWM-A and TRWM-AR. In this case the more recent TRWM-AR was used at the weakness identification stage, first to identify the effects of a failure in a result, using the ASTM E3016-18 (Standard Guide for Establishing Confidence in Digital and Multimedia Evidence Forensic Results by Error Mitigation Analysis) guidance as prompts, and then separately to enumerate the causes or reasons why that effect could manifest.
The results of that analysis are shown below, listing the weaknesses per DFTR (Digital Forensic Technique Result) identified earlier. Each weakness in the knowledge base combines an effect with one cause, and links to its entry in the SOLVE-IT Explorer.
An assertion that an event happened
The weaknesses identified for this result were:
- An AI response to a question does not return a relevant event
- because it found something else relevant and stopped searching (DFW-1422)
- because the data needed to reconstruct the event was lost due to context window overflow (DFW-1433)
- because it was not recognised as relevant (DFW-1444)
- because data needed to reconstruct the event was lost due to compression of the context (DFW-1450)
- An AI response to a question returns an event that did not occur
- An AI response to a question returns an event with one or more incorrect details
- An AI response returns a sequence of events that do not reflect the true ordering of events that occurred
- because of failure to consider clock inaccuracies (DFW-1424)
- An AI response to a question presents an assertion that an event occurred as a fact not a hypothesis (DFW-1425)
- An AI response to a query takes a biased tone
The subset of artefacts deemed relevant
The weaknesses identified for this result were:
- An AI response to a query for relevant artifacts is missing one or more relevant artifacts
- An AI response to a query for relevant artifacts includes reporting artifacts that are not relevant
- because the model could not successfully differentiate relevance (DFW-1432)
- An AI response to a query for relevant artifacts includes reporting artifacts that do not exist
- because of model hallucinations (DFW-1434)
An assertion of a link between entities
The weaknesses identified for this result were:
- An AI response to a question does not return a relevant link
- An AI response to a question asserts a link that does not exist
An explanation of the results presented
The weaknesses identified for this result were:
- An AI response to a query does not provide the reasoning behind the results generated
- An AI response to a query does not provide a measure of confidence in the reasoning provided
- An AI response to a query provides an inaccurate reasoning behind the results generated (DFW-1439)
- An observation referred to in the reasoning provided by an AI system does not exist
- because of model hallucinations (DFW-1440)
- An observation that contradicts the conclusions drawn by an AI system is not presented when presenting reasoning
- An AI response to a query does not provide an accurate confidence in the reasoning provided
- because the confidence value is fabricated (DFW-1443)
- An observation referred to in the reasoning provided by an AI system does not support the conclusions drawn
- because the system is unable to accurately reason with the observable data that is available, but does so anyway (DFW-1445)
Mitigations
So far we have identified 21 mitigations to some of these weaknesses, all of which are integrated into SOLVE-IT. Many weaknesses remain unmitigated; see the section below. Where a mitigation is achieved by applying another SOLVE-IT technique, that technique is shown alongside it.
| Mitigation | Related technique |
|---|---|
| DFM-1332: Flag when an AI context window has reached a limit and prevent further processing within that session | |
| DFM-1343: Use deterministic event reconstruction approaches to complement AI-based event reconstruction | DFT-1086: Analyze timeline |
| DFM-1344: Do not allow context compression during AI-based investigative sessions | |
| DFM-1345: Manually verify a reconstructed event’s existence | |
| DFM-1346: Manually verify the interpretation of artifacts used for an event reconstruction | DFT-1202: Verify automated results manually |
| DFM-1347: Look up in artifact catalogs how the identified artifacts should be interpreted | |
| DFM-1348: Manually verify the properties of a reconstructed event | |
| DFM-1349: Manually verify the artifacts used for an event reconstruction | DFT-1202: Verify automated results manually |
| DFM-1350: Ensure that tools producing data supporting event reconstruction report any potential uncertainty in their results | |
| DFM-1227: Search for indicators of clock tampering | DFT-1129: Search for indicators of clock tampering |
| DFM-1225: Estimate clock offset at a specific point in time using time anchoring | DFT-1134: Use time anchors to estimate clock offset |
| DFM-1333: Use keyword searching to complement AI-based queries for relevant artifacts | DFT-1049: Keyword search |
| DFM-1334: Use timeline analysis to complement AI-based queries for relevant artifacts | DFT-1086: Analyze timeline |
| DFM-1335: Use hash matching to complement AI-based queries for relevant artifacts | DFT-1050: Locate relevant files using a hashset to identify files of interest |
| DFM-1336: Manually verify artifacts reported as relevant by an AI system | DFT-1054: Review artifacts or content manually for relevant material |
| DFM-1337: Manual verification of the existence of artifacts reported by an AI system | DFT-1202: Verify automated results manually |
| DFM-1338: Ensure AI-based tooling provides IDs that link to the artefacts to facilitate manual verification | |
| DFM-1339: Ensure that the AI-based system is capable of providing reasoning behind the selection of results that are relevant | |
| DFM-1340: Ensure that the AI-based system is capable of providing an audit trail of artifacts considered but excluded as relevant, along with the exclusion criteria | |
| DFM-1341: Manually verify an observation used as part of an AI system’s reasoning explanation | DFT-1202: Verify automated results manually |
| DFM-1342: Manually verify the link between entities that is asserted by an AI-based system |
Possible future mitigations?
Despite the work done so far, 18 of the 34 weaknesses listed above have no mitigation in SOLVE-IT yet. Some generalisations are provided below:
- Weaknesses that involve presenting results that are not correct (hallucinations, flawed reasoning etc.) are mitigated through manual checking of results or corroboration with other artefacts or techniques.
- Weaknesses that miss data are mitigated by using other techniques alongside this one, e.g. traditional ‘locate data’ techniques such as keyword searching, timeline analysis, or hash matching. Where the AI system found something relevant and then stopped searching, there is no mitigation yet.
- Weaknesses that are specific to generative AI, such as context window overflow or loss of data through context compression, are mitigated only by features the AI tooling itself has to provide (DFM-1332 and DFM-1344). The knowledge base does not yet link these to any technique that an examiner can apply independently of the tool.
- Presentation issues from the genAI tooling, such as presenting hypotheses as fact, taking a biased tone, failing to disclose uncertainty in results, or stopping prematurely and presenting a concrete conclusion, are not yet addressed in the knowledge base.
This does not mean there are no mitigations for some of these issues, but we have not yet mapped them, and we need help to do so. We also have not enumerated all weaknesses in this new approach, despite the systematic approach used.
If you have weaknesses for this technique and mitigations that can be demonstrated to be effective, please submit them to SOLVE-IT (see below).
Summary
Use of AI agents in digital forensics offers much potential for helping locate relevant digital evidence, particularly in large datasets, in a timely manner. However, once this approach is isolated as a technique in SOLVE-IT, and situated in an investigation with dependencies within a workflow using SOLVE-IT Workflows, it is clear that it adds even more uncertainty to results in an already complex and potentially fragile system. Systematically enumerating weaknesses provides a starting point for managing risk through targeted mitigations, but it needs the collective knowledge of the community to develop this further, and in some cases to create new mitigations for an immature technology bolted on to the digital forensic process.
To submit updates to any SOLVE-IT content you can submit via GitHub directly or use the ‘suggest an edit’ button within the SOLVE-IT Explorer.
-
Hargreaves, C., van Beek, H. and Casey, E., 2025. SOLVE-IT: A proposed digital forensic knowledge base inspired by MITRE ATT&CK. Forensic Science International: Digital Investigation, 52, p.301864. ↩