How can you unlock the potential of AI in your business strategy?
In this article, we will provide you with a framework that will help you tackle this goal systematically.
We will explore this framework through two examples inspired by real experience.
Mark is a compliance officer at a bank working with SMEs—small- and medium-sized enterprises. He is responsible for compliance, reviews customer documents and transactions, and makes decisions about suspicious transactions. Mark is considering AI to analyze documents faster, compare information more accurately and reduce the time required to review transactions.
Anna is an educational programme producer working with students preparing for exams. Together with teachers and curriculum specialists, she is responsible for programme content, practice and learning outcomes. Anna is considering AI to help students study more consistently, receive useful feedback sooner and solve exam questions more successfully on their own.
We will examine both examples using the same framework: from identifying tasks to evaluating AI solutions, business value, risks, costs and integration into the workflow.
Framework for discovering AI potential in your business
Let’s explore a framework for identifying potential AI applications. This framework can uncover fresh perspectives on creating business value with AI.
This framework is inspired by the paper “What Can Machines Learn, and What Does It Mean for Occupations and the Economy?” by Erik Brynjolfsson, Tom Mitchell, and Daniel Rock. Andrew Ng popularised this framework with modern AI.
We will study this framework using the example of the job of a compliance officer at a bank working with SMEs (small- and medium-sized enterprises). Compliance costs for financial institutions are over $206.1 billion per year.
The second example is Anna, an educational programme producer working on exam preparation. Checking students’ work, answering questions and selecting individual practice require substantial team resources. AI could lower these costs, improve student engagement and lead to better learning outcomes.
Analyzing jobs to identify potential AI applications
The basic idea is that AI doesn’t automate full jobs—it automates specific tasks. A job consists of many tasks. Here, “job” can refer to a regular profession (such as Lawyer, HR manager, Merchant, etc.) or to a job in terms of the “jobs to be done” (JTBD) framework.
Here is how you can identify potential AI applications within a job:
- Identify the tasks within a job
- Analyze each task
- Evaluate potential AI solutions
- Evaluate economic benefits
- Evaluate risks
Let’s discuss each step in more detail and see how it applies to a real-life example.
How to identify tasks within a job
Mark — compliance officer at a bank working with SMEs
Consider a bank compliance officer who monitors the transactions of small and medium businesses.
To understand the main tasks a compliance officer is doing it is best to spend a few hours (or days) with several officers to observe their typical activities and ask questions. The expected result of such analysis is a list of the most frequent and important tasks.
Here is a brief overview of the task flow of a compliance officer:
- The officer works with a special web application that provides a queue of alerts about transactions marked as suspicious by the internal automation
- All client information (business profile, uploaded documents, history of transactions, previous alerts, etc.) is available through this web app
- The officer can chat with the client to request additional information
- The officer has access to external sources to check clients’ information
- The officer has to make all decisions strictly based on the specific compliance policies
- The output of the work is the approval or rejection of suspicious transactions and explanatory comments related to these decisions.
The most important tasks of compliance officers:
- Determining the category of customer documents (invoice, rent agreement, permissions, etc)
- Checking for each document:
- Is it real or fake?
- What is the essence?
- How is it related to the customer’s business?
- Who are the counterparties?
- Does it comply with a category-specific policy?
- Checking for each transaction:
- Is it an expected transaction for the type of business?
- Is it a usual amount?
- Does it correspond to the evidence in the customer’s documents?
- Requesting additional documents from the customers if it is impossible to make a decision
- Deciding to approve or reject the transaction
- Talking to the customers to explain the decision about their transactions
- Checking the customer business evidence in external sources
- Tax authorities
- Web presence
- Google reviews
- etc
Anna — educational programme producer
Consider Anna, who organises exam preparation and works with the educational team to help students progress through the programme, practise consistently and achieve their learning goals.
To identify the main tasks, Anna can spend a few hours or days with teachers, curriculum specialists and students: observe study sessions, examine assessed work, review requests for help and ask why particular exercises remain unfinished. The result should be a list of the most frequent and important tasks.
Here is a brief overview of the educational team’s task flow:
- The team works with a learning platform containing materials, exercises, student answers and assessment results.
- For each student, the platform provides the chosen exam, target result, preparation deadline, completed topics, solution attempts and assistance received.
- A teacher or tutor can contact the student, request intermediate steps or clarify what remains unclear.
- Curriculum specialists use the exam specification, assessment criteria and verified learning materials to define programme content.
- Anna works with the team to establish rules for selecting questions, providing feedback and moving to subsequent topics.
- The outputs include a study plan, feedback on solutions, further practice and evidence of the skills the student has mastered independently.
The most important tasks of the educational team:
- Determining the topic and type of question: fractions, percentages, equations, geometry, probability, etc.
- Checking each solution:
- Is there enough evidence to assess the student’s independent understanding?
- What is the essence of the student’s approach?
- How does the solution relate to previously studied skills?
- Which quantities, variables and units are being used?
- Does the solution meet the assessment criteria for the question?
- Checking each practice attempt:
- Is the question appropriate for the student’s current level and preparation goal?
- Is the time taken unusual for this type of question and this student?
- Do the intermediate steps support the final answer?
- Requesting additional working or a clarifying answer when the cause of a mistake is uncertain.
- Deciding whether the student should move to the next topic or complete additional practice.
- Talking to the student to explain the mistake and the recommended next steps.
- Checking learning materials against external sources:
- The official specification for the chosen exam.
- Published assessment criteria.
- Materials verified by teachers.
- Information about changes to exam content.
Anna also connects these tasks to engagement: when students stop studying, whether they return after receiving support and whether they complete another independent attempt. These observations help identify causes of disengagement, but do not replace an assessment of understanding.
How to analyze tasks to identify opportunities for applying AI
Mark — compliance officer at a bank working with SMEs
In the last step, we divided an officer’s work into repeatable tasks. Now we can analyze each to identify opportunities for applying AI.
For each task:
- Evaluate potential AI solutions
- Evaluate economic benefits
- Evaluate risks
You can do these analyses based on your own AI knowledge or in collaboration with your AI team.
When evaluating the potential AI solutions it’s worth using all available sources:
- Papers on relevant topics
- Previous experience of the AI team
- Brainstorm sessions with the team and business stakeholders
- Consultation with domain experts
- Basic analysis of existing data
- Analysis of existing solutions in production
When evaluating the economic benefits of an AI solution, it’s useful to clearly articulate how this solution creates value for the business and its users. Usually, it will be something from one of the following categories:
- Increased revenue
- Lower costs
- Less time to accomplish the task
- Higher accuracy
- Faster processing
- Improved customer service
When evaluating the economic benefits and risks, use rough grading scores (low/medium/high) with some explanations for each choice. Such grading will require judgments from the business stakeholders and will help to align the team.
In our examples we will use a very simple grading:
Low value means that even with a perfect solution, we won’t see any difference in business revenues or costs.
Medium value means that it is worth experimenting with the solutions to understand how AI could influence business metrics.
High value means that the AI solution will definitely accomplish the task more effectively.
Anna — educational programme producer
We have also divided the educational process into repeatable tasks. For each one, Anna works with the AI team, teachers and curriculum specialists to evaluate potential solutions, economic benefits and risks.
The team considers research on feedback and AI-supported learning, its own experiments, examples of educational products in use and actual data on students’ mistakes. Domain experts help determine whether explanations are mathematically correct and useful for learning.
For Anna, the categories of business value have specific meanings:
- Increased revenue: more useful preparation could improve recommendations and encourage students to choose other programmes. This needs to be evaluated in the context of exam seasonality.
- Lower costs: teachers may spend less time on repetitive explanations and preparing for consultations.
- Less time to accomplish the task: students receive support sooner, while teachers get a clearer account of their difficulties.
- Higher accuracy: the team may identify knowledge gaps and select subsequent practice more accurately.
- Faster processing: student answers receive feedback without a long wait.
- Improved customer service: assistance is available at the point of difficulty and takes the current attempt into account.
Anna evaluates learning value separately. More messages to an assistant, longer sessions or more opened lessons do not establish mastery. New questions without hints and a later retention check are needed.
The grades in the educational case are preliminary. For example, understanding a mistake has medium expected value: the task affects preparation quality, but the additional effect of AI needs measurement. The consequences of an incorrect hint are assessed separately from the likelihood that a student will recognise the error.
Let’s examine two examples:
- What is the essence of a document
- Determining the category of customer documents
Alongside Mark’s tasks, we will examine two tasks from Anna’s educational process:
- Understanding the essence of a student’s solution and the reason for a mistake.
- Determining the topic and type of a question.
Task: Understanding documents and students’ solutions
Mark — understanding the essence of a document
Clients can provide many types of documents: invoices, rent agreements, certificates, business contracts, licenses, etc.
For each transaction, the officer must download and open all documents one by one, look through their content, and analyze the important details.
For example, for an invoice, it’s important to take into account:
- Counterparties
- Date of the invoice and due date
- Amount
- Purpose of the invoice
- Terms of payment
Based on these fields, the officer decides if it is a valid invoice using the bank compliance policy. For instance, the invoice must be related to the bank client, the purpose of the invoice should correspond to the nature of the client’s business, it should have the correct dates and currency, etc.
This task is very time-consuming because documents could be dozens of pages, they can have poor quality (for example, it could be just a screenshot of a handwritten invoice), and there could be a lot of documents per client.
How AI can be applied
- OCR (optical character recognition) can be used to extract text from scanned documents
- LLMs can be used to extract key information, interpret the context of the document, summarize it, and answer questions based on the document’s content.
Business value
- Less time to accomplish the task: AI reduces the manual effort required to understand the document.
- Higher accuracy: By accurately extracting and summarizing key information, AI supports more informed decisions.
- Medium value:
Understanding long documents take a significant share of an officer’s time. For example, a 20-page contract could require around five minutes. AI can reduce this time to less one minute by providing a structured summary of the document.
Severity of risk
Medium risk
If the AI misinterprets nuanced or context-dependent information, it could lead to incorrect conclusions or actions by the officer.
Anna — understanding the essence of a student’s solution
Students can submit different types of work: short answers, detailed calculations, written explanations, photographs of handwritten solutions and graphs.
For each attempt, the teacher needs to open the question, examine the working, compare it with valid methods and understand the important details.
For example, when reviewing an equation, it is important to identify:
- Which equation the student is solving.
- Which method they have selected.
- Which transformations they have made and in what order.
- The step at which a discrepancy appears.
- Whether the problem is a calculation slip or a misunderstanding of a rule.
- Whether the final answer is supported by the working.
Using this information and the assessment criteria, the teacher determines what support is needed. A correct solution using an alternative method should be accepted, while a correct final answer following invalid transformations should not automatically count as evidence of understanding.
This task can consume a significant share of the team’s time. Solutions may be incomplete, handwriting and mathematical notation ambiguous, and apparently similar mistakes may have different causes. As student numbers increase, individual feedback can be delayed.
Consider this example: 3(x + 2) = 15 → 3x + 2 = 15. One possible cause is a misunderstanding that the multiplier applies to both terms. However, a single line does not rule out a careless slip. The assistant could ask, “How would you expand 2(a + 4)?” and use the response to choose an explanation.
How AI can be applied
- OCR and image recognition can be used to extract text and mathematical working from photographs. Formula recognition needs its own quality evaluation.
- LLMs can structure the steps, compare them with a verified explanation, suggest a possible cause of the mistake and prepare feedback.
- Verified answers and computational tools can check supported transformations. When a solution is ambiguous, the system should request an additional step or refer the case to a teacher.
- The history of previous attempts can help select a clarifying question and avoid repeating an explanation the student has already found unhelpful.
Business value
- Less time to accomplish the task: AI could reduce manual effort in initial analysis and feedback preparation.
- Higher accuracy: a structured account of the working could help a teacher distinguish a recurring gap from an isolated slip. This benefit needs testing.
- Improved engagement: timely assistance could increase the likelihood that the student finishes an exercise after making a mistake.
- Medium value: the task meaningfully affects the student experience and teacher workload, but the effect of AI on support costs and independent performance requires an experiment. Anna compares time to useful feedback, substantive teacher corrections and performance on the next independent attempt.
Severity of risk
- Medium risk in a limited scenario with verified questions and quality checks: AI may misread a symbol, reject a valid method or incorrectly identify the cause of a mistake. A student could learn an incorrect rule. Using the same outputs to assign a final grade or determine exam readiness would carry a higher risk.
Task: Determining document categories and question topics
Mark — determining the category of customer documents
When a client uploads a new document, the officer must determine its category (invoice, rent agreement, license, etc.).
The category will be used to guide the verification process because each category has its own compliance policy. Category information is entered through a special web interface and stored for further use.
How AI can be applied
- Natural language processing (NLP), LLMs, OCR, and image recognition technologies can be used to automatically classify documents based on their content.
Business value
- Less time to accomplish the task: AI reduces the time spent reading, sorting, and processing the documents.
- Low value
Officers can do the classification almost instantly. Basic automated classification is already in place.
Severity of risk
Low risk
Incorrect classification can be fixed during manual review
Anna — determining the topic and type of a question
When a new exercise is added to the educational project, a curriculum specialist determines its topic and type. Examples include operations with fractions, linear equations, percentages, trigonometry and probability.
The category is used to select practice and analyse results: the question is associated with specific skills, a preparation level and a part of the exam specification. This information is stored in the question bank and used in subsequent recommendations.
How AI can be applied
- Natural language processing, LLMs, OCR and image recognition can analyse the wording, formulae and diagrams and suggest a topic.
- The system can suggest several skill labels for a multi-step question, which a curriculum specialist then confirms.
Business value
- Less time to accomplish the task: AI could reduce manual labelling of new exercises.
- Low value in this project: the main question bank is already labelled, and curriculum specialists can quickly handle the current volume of additions. Even accurate classification would not yet produce a meaningful change in costs or student outcomes.
Severity of risk
- Low risk with manual review: a curriculum specialist can correct the category before the question is used. Without review, incorrect labels could lead to unsuitable practice and inaccurate conclusions about a student’s knowledge gaps.
To improve your skills in such analysis, study our AI/ML Simulator for Product Managers.
Analysis of all tasks
Mark — all the main tasks of a compliance officer
Here is a summary table for all the main tasks of a compliance officer:
| Task | How AI can be applied | Business value | Severity of risk |
|---|---|---|---|
| Determining the category of customer documents | Natural language processing (NLP), LLMs, OCR and image recognition technologies can be used to automatically classify documents based on their content | Less time to accomplish the task: AI reduces the time spent reading, sorting, and processing the documents. Low value: Officers can do the classification almost instantly. Basic automated classification is already in place. | Low risk:Incorrect classification can be fixed during manual review |
| Is the document real or fake? | Computer vision models can analyze document features to detect anomalies or signs of forgery that might indicate falsification. | Higher accuracy: enhance forgery detection. Low value:Forgeries are very rare, manual processing is enough. | Low risk: False positives can be mitigated during manual review |
| What is the essence of a document? | OCR can be used to extract text from scanned. LLMs can be used to extract key information, interpret the context of the document, summarize it, and answer questions based on its content. | Less time to accomplish the task: AI reduces the manual effort required to understand the document. Better accuracy: By accurately extracting and summarizing key information, AI supports more informed decisions. Medium value: Understanding long documents takes a lot of time. | Medium risk: If the AI misinterprets nuanced or context-dependent information, it could lead to incorrect conclusions or actions by the officer. |
| How is the document related to the customer’s business? | AI systems can compare a document details to business profiles to assess its relevance. | Higher accuracy: Improving the relevance of compliance checks. Medium value: Understanding long documents takes a lot of time. | Medium risk: AI could mistakenly associate documents with incorrect business contexts, leading to inappropriate compliance actions. |
| Identifying the counterparties in the document | LLMs can extract named entities and their roles within documents. | Less time to accomplish the task: AI reduces the need to manually read documents and fill forms. Low value: It’s easy for officers to do this task manually. | Medium risk: Errors in entity recognition could misidentify or omit critical parties involved, affecting compliance integrity. |
| Does the document comply with a category-specific policy? | Rule-based AI systems can verify whether documents adhere to specific regulatory frameworks and industry standards. | Less time to accomplish the task: AI reduces the manual effort required to analyze the document. Better accuracy: reducing human errors in complex policy criteria evaluation. High value: Making decisions about a document can take a lot of time. | High risk: Incorrect policy interpretations could lead to regulatory risks for the business. |
| Is this transaction expected for this type of business? | ML models can compare transactions to typical business activities. | Less time to accomplish the task: AI reduces manual document lookups and speeds up the compliance workflow. Medium value: Understanding the business profile is a time-consuming activity. | Medium risk: As AI relies on historical data, it may not predict unconventional-yet-legitimate business practices, potentially flagging them as suspicious, leading to inappropriate compliance actions. |
| Is the transaction amount normal for this kind of business? | Algorithms can analyze historical data to flag transactions that deviate from established norms. | Higher accuracy: Predictive ML models can flag unusual transaction amounts. Low value: Basic automated checks are already in place. | Medium risk:AI might flag normal transactions as suspicious if they deviate from typical patterns, leading to unnecessary investigations. |
| Does the transaction correspond to the evidence in customer documents | Transaction details can be automatically matched with corresponding customer documents using the results of a named entity extraction system. | Less time to accomplish the task: AI reduces manual document lookups and speeds up the compliance workflow. Medium value: Understanding the business profile is time-consuming. | Medium risk: AI could misassociate transactions and customer documents leading to inappropriate compliance actions. |
| Deciding to approve or reject the transaction | Automated approvals (rule-based system) can be based on predefined criteria. | Faster processing: AI can automate routine decision-making, speeding up transaction approvals and reducing bottlenecks. Medium value: A basic rule-based system is already in place. | High risk: Over-reliance on AI decision-making can reduce human oversight and increase the risk of errors. Customer satisfaction could deteriorate. |
| Requesting additional documents from the customer | Automated systems can trigger requests for additional documentation given the analysis of existing data. | Faster processing: Automated document requests streamline interactions and ensure necessary documentation is collected without delay. High value: The faster the process goes for users, the higher their satisfaction. | Medium risk: Automated requests could be triggered inappropriately, leading to customer dissatisfaction or data overload. |
| Talking to customers to explain the decision about their transactions | LLM-powered chatbots can handle routine inquiries and provide explanations regarding transaction decisions. | Improved customer service: AI-driven chatbots can provide instant responses to customer queries, improving satisfaction and engagement. High value: The faster the process goes for users, the higher their satisfaction. | High risk: AI-generated responses may contain hallucinations, incorrect information, leak personal data, lack empathy, or fail to address specific customer concerns adequately, potentially harming customer relations and causing legal risks for the business. |
| Checking the customer business evidence with external data | External databases and APIs (like tax records) can provide data to verify and enhance customer information. LLMs can summarize content of web pages. | Lower costs: Automating the verification process with AI reduces the need for extensive manual background checks, cutting operational costs. Medium value: The manual process consumes a lot of time. | Low risk: AI could rely on outdated or incorrect external data, leading to inaccurate assessments of compliance status. |
Anna — all the main tasks of the educational team
The following table provides the parallel educational analysis. The ratings apply to the proposed scenario and need to be tested using project data.
| Task | How AI can be applied | Business value | Severity of risk |
|---|---|---|---|
| Determining the topic and type of a question | NLP, LLMs, OCR and image recognition can analyse wording, formulae and diagrams and suggest categories and skills. A curriculum specialist confirms the labels. | Less time spent labelling. Low value: the bank is already labelled and the current volume of additions is small. | Low risk with manual review: incorrect labels can be fixed before use. Unchecked errors affect recommendations. |
| Checking whether the work demonstrates independent understanding | AI can compare explanations with recorded steps and suggest a short follow-up question or an oral explanation. Writing style alone cannot establish that a student copied a solution. | Better-supported diagnosis. Medium value: additional checks can be useful, but their effect on workload and learning needs measurement. | High risk if used for automatic accusations or grading: independent work could be wrongly treated as copied. The pilot uses follow-up questions rather than accusations. |
| Understanding the essence of a solution and the reason for a mistake | OCR can extract the working, while LLMs can structure the steps, compare them with verified methods and propose a hypothesis about the difficulty. | Less time spent on initial analysis and feedback preparation. Medium value: individual review consumes substantial time, but the benefit of AI needs confirmation. | Medium risk with quality checks: a misread symbol or incorrect interpretation can lead to an unsuitable explanation. |
| Understanding how a mistake relates to previously studied skills | The system can compare an attempt with earlier work and a map of skill dependencies. For example, it can investigate whether a difficulty with equations relates to operations with fractions. | More targeted revision. Medium value: the team could identify prerequisite gaps sooner and avoid unnecessary revision of an entire topic. | Medium risk: an incorrect connection could send the student to an unsuitable topic or leave a necessary prerequisite unaddressed. |
| Identifying quantities, variables and units | LLMs can extract notation from the question and the solution and compare its roles. | Less manual effort during review. Low value: a teacher or an existing template can quickly identify these elements in routine questions. | Medium risk: an ambiguous symbol or missing unit can distort a subsequent explanation. |
| Checking whether a solution meets assessment criteria | The system can compare recorded steps with the verified mark scheme for a particular question and prepare a provisional explanation of marks for a teacher. | Faster feedback preparation for extended work. Potentially high value at a large assessment volume; actual savings need confirmation. | High risk: the model could misapply criteria or overlook a valid method. A provisional analysis should not automatically become a final grade. |
| Determining whether a question suits the student’s level and goal | Recommendation models can use the initial diagnosis, recent attempts, chosen exam, level and available time to select questions from a verified bank. | More suitable practice. Medium value: potential improvements in consistent practice and skill mastery need comparison with a standard study plan. | Medium risk: overly easy practice creates an illusion of progress, while overly difficult practice may cause students to stop. |
| Determining whether the time taken is unusual | Algorithms can compare the duration of an attempt with the student’s previous attempts and the question’s characteristics. | Faster identification of possible difficulties. Low value if simple rules are sufficient and the platform already records inactivity. | Medium risk: a long attempt may reflect a break, while a short one may reflect a familiar question. Time alone does not establish understanding. |
| Checking whether the final answer corresponds to the working | LLMs and computational tools can check supported transformations, substitute results and examine whether the conclusion follows from the recorded steps. | More informative assessment than comparing the final answer alone. Medium value: it can reveal accidentally correct results and guide feedback. | Medium risk: valid shortcuts and alternative methods could be wrongly flagged as inconsistent. Ambiguous cases require clarification. |
| Deciding whether to move to the next topic | The system can summarise independent attempts and later checks and recommend progression or additional practice. | Faster study plan updates. Medium value: benefits need comparison with rules based on assessment questions. | High risk if based on one answer or attempts with hints: an undetected gap could undermine subsequent learning. |
| Requesting additional working | LLMs can ask a focused question about a missing step or suggest a short diagnostic attempt. | Faster collection of information about the difficulty and less teacher correspondence. Medium value: the outcome depends on whether requests are relevant and understandable. | Medium risk: unnecessary questions can frustrate students, while an overly revealing prompt can distort diagnosis. |
| Talking to students to explain a mistake and the next action | LLMs can prepare explanations using verified materials, the current attempt and support already provided. | More accessible feedback and an opportunity to continue studying. Potentially high value at a large support volume, but independent learning effects need testing. | High risk without appropriate limits: a confident incorrect explanation can reinforce a misconception, while a complete answer can replace independent practice. |
| Checking materials against external exam information | The system can compare materials with the selected official specification and assessment criteria, record the source and version and prepare discrepancies for a curriculum specialist. | Less time spent checking and updating materials. Medium value when supporting several exams and regular content updates. | Medium risk: an outdated specification, different tier or different exam board can lead to inappropriate preparation. A curriculum specialist reviews changes. |
Choosing the scope of the project
Deciding on how and where to apply AI differs case by case.
Consider some basic criteria:
- Solving the task should bring business value (decrease costs, increase revenue, etc.)
- AI solution should be feasible
- The risks of applied AI should not be too high
- The costs of the solution should be less than the business benefits
Mark and Anna apply the same criteria to their tasks.
The detailed scope decisions for Mark’s original case appear below. Anna’s initial area of investigation is understanding a solution, clarifying the cause of a mistake and selecting subsequent practice. Conclusions about exam readiness have different consequences and require a separate quality evaluation.
As a trial example for the educational implementation, consider passanexam.com, a mathematics exam preparation project, specifically for GCSE Maths. Its website presents an AI tutor called Ari, with a level check, a personal plan, explanations and practice with feedback.
In our hypothetical pilot, Anna selects a few algebra skills within one exam and tier: for example, expanding brackets, manipulating expressions and solving linear equations. Each topic needs verified questions, valid methods, examples of mistakes and independent assessment questions. This is a proposed experiment, rather than a description of confirmed Pass an Exam results.
Evaluating the feasibility of the AI solutions
The solution can likely be delivered if:
- Your AI team successfully solved a similar problem
- Other teams of competitors successfully solved a similar problem with AI
- There are plenty of papers on the topic
- There are cloud services that provide a similar solution
- There are open-source libraries that provide a similar solution
- All the necessary data is available
In some cases, the team could face difficulties. It will be best to start with simple and quick prototypes to gain a better understanding of the complexity, costs, and risks.
If external tools/services/libraries are available, it is worth using them to quickly get a feel of the potential business value.
Applying this to Mark.
In the original case, the candidate tasks involve processing documents and comparing information. Mark and the AI team can test selected tools on documents approved for evaluation, including long documents, poor scans and ambiguous cases. For each output, they need to check consequential fields and the supporting source passage. An available tool makes prototyping easier, but performance on the bank’s own documents needs evaluation.
Applying this to Anna.
Anna examines the same grounds for feasibility in mathematics preparation:
- Does the AI team have experience analysing mathematical working, selecting practice or preparing educational feedback?
- Are there examples of working solutions for a comparable exam, question type and student age group?
- Which studies and experiments help determine the hint format and the method for checking independent mastery?
- Which available models and services can work with text, formulae and images?
- Which libraries can check calculations and supported transformations?
- Are verified questions, valid solutions, assessment criteria and examples of actual student mistakes available?
For the pilot using passanexam.com as an example, Anna proposes starting with step-by-step text input. Photographs introduce a separate recognition task and another source of error, so the team can evaluate them in a subsequent stage.
Teachers prepare a collection containing valid alternative methods, common mistakes, incomplete working and cases with insufficient information. The team evaluates mathematical accuracy, clarity, the relevance of follow-up questions and whether students can continue. A simple prototype should also reveal how much human review the actual workflow would require.
Analyzing costs
To estimate the costs of an AI solution, consider the following:
- Compute costs for using large GenAI models or cloud services
- Efforts of the AI team
Evaluate compute costs per unit of work. In our case, it could be the cost of processing one document or transaction. Such evaluation will allow us to compare the time and costs required for manual and automated processing.
Team efforts will depend on the specifics of the subject area, the qualifications of the team, and the availability of off-the-shelf solutions. Rely on the team’s experience: if they have previously solved similar problems, expectations about the results will be more realistic.
For Mark.
The unit of work remains a document or transaction. The team also includes the officer’s time spent checking AI outputs and correcting mistakes: generating a summary quickly does not save time if it requires extensive rechecking. Implementation and maintenance costs need to be allocated across a realistic processing volume.
For Anna.
The unit of work could be the analysis of one solution, a study session with several hints or a month of support for an active student. For each unit, the team evaluates:
- The number of model calls and the amount of context supplied.
- Recognition and computational checking costs, where applicable.
- The effort required to prepare and update verified questions and explanations.
- Teacher time for quality checks and uncertain cases.
- Integration, analytics, maintenance and error-correction effort.
The cost of a session includes both automated processing and human assistance. Anna compares it with current support costs at comparable quality and also examines the cost of achieving a verified learning outcome.
More frequent use of the assistant can increase AI costs. If teachers must correct a substantial share of the explanations, the expected savings may disappear. The pilot therefore needs to evaluate the total cost of useful feedback and the following independent attempt, alongside the price of an individual model call.
Let’s apply this logic to our examples.
Task: Understanding documents and students’ solutions
Mark — understanding the essence of a document
- Accomplishing the task should bring business value itself regardless of AI: medium value is expected – it takes less time to accomplish the task and it will provide higher accuracy.
- AI solutions are feasible and can bring business improvements in comparison to the status quo: it’s easy to find lots of examples of similar solutions (named entity recognition, summarization, OCR).
- The risks of applied AI are not too high.
- Costs of the solution should be reasonable: compute costs can be very low when using cloud services or hosted models.
Summary for this task: we should include it in the scope of the project.
Anna — understanding the essence of a student’s solution
- Accomplishing the task should bring business value regardless of AI: medium value is expected. The team could analyse difficulties and prepare support sooner, but changes in teacher workload and independent performance need measurement.
- AI solutions should be feasible and improve on the current process: a limited set of questions can be used to test step analysis, clarifying questions and explanations based on verified materials. A working prototype does not itself establish adequate quality.
- The risks of applied AI should be manageable: the pilot uses verified questions, teacher evaluation of hints and referral of ambiguous cases to a human. An attempt completed with a hint is not treated as sufficient evidence of mastery.
- Costs should be reasonable: model usage, content preparation and quality checks are compared with the current cost of feedback and the learning effect achieved.
Summary for this task: include it in the pilot’s scope. Decide whether to expand after evaluating quality, engagement and performance on independent questions.
Task: Determining document categories and question topics
Mark — determining the category of customer documents
- Solving the task should bring business value regardless of AI: low value – officers can do the classification almost instantly. Basic automated classification is already in place.
- AI solutions are feasible and can bring business improvements in comparison to the status quo: it’s easy to find many examples of similar solutions (document classification).
- The risks of applied AI are low.
- Costs of the solution should be reasonable: compute costs can be very low when using cloud services or hosted models.
Summary for this task: we should not include it in the scope of the project.
Anna — determining the topic and type of a question
- Solving the task should bring business value regardless of AI. Low value is expected: the existing bank is labelled and the current volume of additions does not create a substantial workload.
- An AI solution is feasible: automatic classification can be tested against existing labels, including questions that assess several skills.
- The risks of applied AI are low when a curriculum specialist checks the labels before publication.
- Costs should be reasonable: even with low compute costs, integration, corrections and classifier maintenance must be included.
Summary for this task: do not include it in the initial pilot. Reconsider when the volume of new questions makes manual labelling a significant cost.
Appropriate tasks for the project
Mark — the original project’s tasks
Based on the logic described above we can finalize the scope of the project:
- What is the essence of a document
- How is a document related to the customer’s business
- Is this transaction expected for this type of business
- Does the transaction correspond to the evidence in customer documents
- Requesting additional documents from the customers
- Checking the customer business evidence
Anna — the educational pilot’s tasks
Based on the analysis, Anna selects:
- Understanding the essence of a student’s solution and the possible reason for a mistake.
- Understanding how the difficulty relates to previously studied skills.
- Determining whether the next question suits the student’s current level and preparation goal.
- Checking whether the final answer corresponds to the recorded working for supported question types.
- Requesting additional steps or a short diagnostic attempt when information is insufficient.
- Preparing explanations from verified materials, with quality checks and access to teacher support.
- Checking materials against the selected exam specification under a curriculum specialist’s supervision.
In the GCSE Maths pilot using passanexam.com as an example, these tasks form a learning sequence: independent attempt → analysis of the difficulty → clarification if needed → short hint → new independent question → later skill check.
Progression to the next topic relies on assessment questions and rules agreed with curriculum specialists. An automatic conclusion about exam readiness is outside this limited pilot’s scope.
What’s next
Mark — integration into the compliance officer’s workflow
We’ve chosen the tasks of the compliance officer where AI is most promising. The next step should be to outline the architecture of the AI solution. This is a critical point – now we have to go back to the officers and determine how the solution will be integrated into the current web interface to gain the highest business value.
It’s necessary to understand
- How the officer will access the results of the AI system
- How the results will be presented in the interface
- How the officer’s feedback about the AI-generated results will be incorporated into the workflow
It’s best to address these questions before developing the AI solution because the desired UX is crucial and can shape the underlying solution.
For instance:
- If the summary of the document is too long and unstructured, then it won’t save time for the officer. It’s better to talk to the officers to determine the best format for document summaries.
- If officers have to wait for the AI models to process the documents every time they want to access the extracted information, then they will lose precious time. It’s better to process the documents in the background as soon as they are uploaded by the customers.
- If the officer needs to use some unfamiliar external system (for example Jupyter Notebook) to access the extracted information, then it could undermine their motivation to use the AI. It’s better to integrate the AI-generated results directly into the web application that the officers use.
- If the officer is unable to flag errors in the AI-generated results or dispute them, or if the officer can’t understand the reasoning behind the AI’s decision-making, then they won’t trust the AI solution and the whole project might fail.
Anna — integration into the student’s preparation
We have selected educational tasks where AI looks promising. The next step is to outline the architecture and return to students, teachers and curriculum specialists to determine how the support should fit into the platform and study sessions.
It is necessary to understand:
- How students will access support in the context of a particular question, and how teachers will access information about difficulties.
- How the hint, explanation, step under review and subsequent practice will appear in the interface.
- How the system will distinguish independent attempts, solutions completed with hints and later assessments.
- How reports of unclear explanations, teacher corrections and subsequent attempts will be collected.
These questions should be addressed before development: the interaction determines which data to collect, when to call the model and what it should return.
For instance:
- If an explanation is long and contains the complete solution, a student may read it and move on without making another attempt. Anna should test short hints, retrying the current step and progressively revealing assistance with teachers and students.
- If students wait a long time for support after every error, they may stop studying. Verified materials and question context can be prepared in advance, with personalised feedback generated for the actual attempt. Acceptable response times need testing in the prototype.
- If students must switch to a separate chat and copy the question and their working again, some context will be lost. In the proposed passanexam.com pilot, assistance should be connected to the current question and relevant step within exam preparation.
- If students cannot flag incorrect or unclear feedback, and teachers cannot see the original attempt and correct the analysis, the team cannot reliably improve quality. A clear feedback process and a way to review uncertain cases are needed.
- If the platform marks a skill as mastered immediately after a solution with hints, the subsequent plan may rely on an inflated assessment. New questions and later checks should therefore be completed without assistance, with the different attempt types recorded separately.
Anna defines several groups of pilot measures:
- Engagement: return to practice, exercise completion after an error and completion of planned sessions.
- Learning outcomes: performance on new questions without hints, recurrence of earlier mistakes, retention after several days and performance on a comparable assessment.
- Support quality: incorrect or inappropriate explanations, substantive teacher corrections and referrals to a human.
- Economics: session cost, teacher support time and cost per active student.
Anna plans to compare groups receiving the same materials and practice, with one group also receiving the AI support under investigation. Assignment would be randomised with initial attainment taken into account. Final questions would be new, comparable in difficulty and completed without assistance. The team determines sample size and experiment duration based on the selected metric and the effect it needs to detect.
A small preliminary launch can reveal errors and usability problems, but cannot by itself establish improved exam results. If interaction with AI increases while independent performance remains unchanged, the support format needs revision. If skill mastery and consistent practice improve at an acceptable cost, the team can expand the topics covered.
Summary
To discover potential AI applications, you can use a simple framework:
- Identify the tasks within a job
- Analyze each task
- Evaluate the potential AI solutions
- Evaluate the economic benefits
- Evaluate the risks
To understand the main tasks of a job it is best to spend a few hours (or days) with professionals who do the job to observe their typical activities and ask questions.
When evaluating potential AI solutions, it’s worth using all available sources like papers on relevant topics, previous experience of the AI team, brainstorming sessions with the team, etc.
When evaluating the economic benefits it’s useful to clearly articulate how this solution creates value for the business and users (increased revenue, lower costs, less time to accomplish the task, etc.).
When evaluating the economic benefits and risks of AI solutions, use rough grades low/medium/high with a brief explanation of the choice.
Basic criteria for choosing the scope of a project:
- Solving the task should bring business value
- The AI solution should be feasible
- The risks of applied AI should not be too high
- The costs of the solution should be less than the business benefits
Before outlining the architecture of the AI solution, it’s necessary to understand:
- How the results of AI automation will be accessed
- How they will be presented to the users
- How user feedback about the AI-generated results will be processed
How the framework connects Mark’s and Anna’s examples.
Mark applies it to compliance work: examining documents, comparing information and preparing the evidence needed for transaction decisions. The tasks, ratings and conclusions from the original example have been preserved above.
Anna applies the same sequence to the educational process: analysing a solution, identifying a difficulty, selecting support and assigning subsequent practice. For each task, she evaluates potential AI solutions, business value, risks, feasibility and costs.
The trial example of passanexam.com and GCSE Maths gives the educational pilot a specific exam, tier and set of skills. Feedback quality, consistent practice, independent performance and acceptable support costs determine whether to expand. The expected effect needs experimental confirmation.
In both cases, the technology’s usefulness is reflected in the work’s outcome: the quality and speed of Mark’s reviews, and how successfully Anna’s students learn and prepare for their exams.
Learn more
To train your AI product skills, try our simulators:
Illustration by Anna Golde for GoPractice