


What KPIs can be used to measure the success of artificial intelligence projects?
A research report released by the research firm IDC in June 2020 showed that approximately 28% of artificial intelligence plans failed. Reasons cited in the report were a lack of expertise, a lack of relevant data and a lack of a sufficiently integrated development environment. In order to establish a process for continuous improvement of machine learning and avoid getting stuck, identifying key performance indicators (KPIs) is now a priority.
#In the upper reaches of the industry, data scientists can define the technical performance indicators of the model. They will vary depending on the type of algorithm used. In the case of a regression aimed at predicting someone's height as a function of their age, for example, one can resort to linear determination coefficients.
An equation to measure the quality of the prediction can be used: If the square of the correlation coefficient is zero, the regression line determines the 0% point distribution. On the other hand, if the coefficient is 100%, the number is equal to 1. Therefore, this indicates that the quality of the predictions is very good.
Deviation of predictions from reality
Another metric for evaluating regression is the least squares method, which refers to the loss function. It involves quantifying the error by calculating the sum of squared deviations between the actual value and the predicted line, and then fitting the model by minimizing the squared error. In the same logic, one can utilize the mean absolute error method, which consists in calculating the average of the fundamental values of the deviations.
Charlotte Pierron-Perlès, who is responsible for strategy, data and artificial intelligence services at French consultancy Capgemini, concluded: "In any case, this amounts to measuring the gap with what we are trying to predict."
For example, In classification algorithms for spam detection, it is necessary to look for false positives and false negatives of spam. Pierron Perlès explains: "For example, we developed a machine learning solution for a cosmetics group that optimizes the efficiency of a production line. The aim was to identify defective cosmetics at the beginning of the production line that could cause production interruptions. We worked closely with the factory operators Discussions followed with them seeking a model to complete the detection even if it meant detecting false positives, that is, qualified cosmetics could be mistaken for defective."
Based on false positives and false negatives The concept of three other metrics allows the evaluation of classification models:
(1) Recall (R) refers to a measure of model sensitivity. It is the ratio of correctly identified true positives (taking positive coronavirus tests as an example) to all true positives that should have been detected (positive coronavirus tests and negative coronavirus tests were actually positive): R = true positives / true positives false Negative.
(2) Precision (P) refers to the measure of accuracy. It is the ratio of correct true positives (positive COVID-19 tests) to all results determined to be positive (positive COVID-19 tests negative COVID-19 tests): P = true positives / true positives false positives.
(3) Harmonic mean (F-score) measures the model’s ability to give correct predictions and reject other predictions: F=2×precision×recall/precision-recall
of the model Promotion
DavidTsangHinSun, chief senior data scientist at French ESNKeyrus company, emphasized: "Once a model is built, its generalization ability will become a key indicator."
So how to estimate it? By measuring the difference between predictions and expected results, and then understanding how that difference evolves over time. He explains, "After a period of time, we may encounter divergence. This may be due to underlearning (or overfitting) due to insufficient training of the data set in terms of quality and quantity."
So what is the solution? For example, in the case of image recognition models, adversarial generative networks can be used to increase the number of pictures learned through rotation or distortion. Another technique (applicable to classification algorithms): synthetic minority oversampling, which consists of increasing the number of low-occurrence examples in the data set through oversampling.
Disagreement can also occur in the case of over-learning. In this configuration, the model will not be restricted to the expected correlations after training, but due to overspecialization, it will capture the noise generated by the field data and produce inconsistent results. DavidTsangHinSun pointed out, "It is then necessary to check the quality of the training data set and possibly adjust the weight of the variables."
While the economic key performance indicators (KPIs) remain. Stéphane Roder, CEO of French consulting firm AIBuilders, believes: “We have to ask ourselves whether the error rate is consistent with the business challenges. For example, the insurance company Lemonade has developed a machine learning module that can respond to customer requests within 3 minutes after filing a claim. information (including photos) to pay insurance benefits to the customer. Taking into account the savings, a certain error rate incurs costs. Over the entire life cycle of the model, especially compared to the total cost of ownership (TCO), from development to maintenance , it is very important to check this measurement value."
Adoption Level
Even within the same company, expected key performance indicators (KPIs) may vary. Charlotte Pierron Perlès of Capgemini noted: "We developed a consumption forecasting engine for a French retailer with an international standing. It turned out that the precise targeting of the model differed between products sold in department stores and new products. Sales of the latter Dynamics depend on factors, especially those related to market reaction, which are, by definition, less controllable."
The final key performance indicator is adoption levels. Charlotte Pierron-Perlès said: "Even if a model is of good quality, it is not enough on its own. This requires the development of artificial intelligence products with a user-oriented experience that can be used for business and realize the promise of machine learning."
Stéphane Roder concluded: “This user experience will also allow users to provide feedback, which will help provide artificial intelligence knowledge outside the daily production data flow.”
The above is the detailed content of What KPIs can be used to measure the success of artificial intelligence projects?. For more information, please follow other related articles on the PHP Chinese website!

Hot AI Tools

Undresser.AI Undress
AI-powered app for creating realistic nude photos

AI Clothes Remover
Online AI tool for removing clothes from photos.

Undress AI Tool
Undress images for free

Clothoff.io
AI clothes remover

AI Hentai Generator
Generate AI Hentai for free.

Hot Article

Hot Tools

Notepad++7.3.1
Easy-to-use and free code editor

SublimeText3 Chinese version
Chinese version, very easy to use

Zend Studio 13.0.1
Powerful PHP integrated development environment

Dreamweaver CS6
Visual web development tools

SublimeText3 Mac version
God-level code editing software (SublimeText3)

Hot Topics



This site reported on June 27 that Jianying is a video editing software developed by FaceMeng Technology, a subsidiary of ByteDance. It relies on the Douyin platform and basically produces short video content for users of the platform. It is compatible with iOS, Android, and Windows. , MacOS and other operating systems. Jianying officially announced the upgrade of its membership system and launched a new SVIP, which includes a variety of AI black technologies, such as intelligent translation, intelligent highlighting, intelligent packaging, digital human synthesis, etc. In terms of price, the monthly fee for clipping SVIP is 79 yuan, the annual fee is 599 yuan (note on this site: equivalent to 49.9 yuan per month), the continuous monthly subscription is 59 yuan per month, and the continuous annual subscription is 499 yuan per year (equivalent to 41.6 yuan per month) . In addition, the cut official also stated that in order to improve the user experience, those who have subscribed to the original VIP

Improve developer productivity, efficiency, and accuracy by incorporating retrieval-enhanced generation and semantic memory into AI coding assistants. Translated from EnhancingAICodingAssistantswithContextUsingRAGandSEM-RAG, author JanakiramMSV. While basic AI programming assistants are naturally helpful, they often fail to provide the most relevant and correct code suggestions because they rely on a general understanding of the software language and the most common patterns of writing software. The code generated by these coding assistants is suitable for solving the problems they are responsible for solving, but often does not conform to the coding standards, conventions and styles of the individual teams. This often results in suggestions that need to be modified or refined in order for the code to be accepted into the application

To learn more about AIGC, please visit: 51CTOAI.x Community https://www.51cto.com/aigc/Translator|Jingyan Reviewer|Chonglou is different from the traditional question bank that can be seen everywhere on the Internet. These questions It requires thinking outside the box. Large Language Models (LLMs) are increasingly important in the fields of data science, generative artificial intelligence (GenAI), and artificial intelligence. These complex algorithms enhance human skills and drive efficiency and innovation in many industries, becoming the key for companies to remain competitive. LLM has a wide range of applications. It can be used in fields such as natural language processing, text generation, speech recognition and recommendation systems. By learning from large amounts of data, LLM is able to generate text

Large Language Models (LLMs) are trained on huge text databases, where they acquire large amounts of real-world knowledge. This knowledge is embedded into their parameters and can then be used when needed. The knowledge of these models is "reified" at the end of training. At the end of pre-training, the model actually stops learning. Align or fine-tune the model to learn how to leverage this knowledge and respond more naturally to user questions. But sometimes model knowledge is not enough, and although the model can access external content through RAG, it is considered beneficial to adapt the model to new domains through fine-tuning. This fine-tuning is performed using input from human annotators or other LLM creations, where the model encounters additional real-world knowledge and integrates it

Editor |ScienceAI Question Answering (QA) data set plays a vital role in promoting natural language processing (NLP) research. High-quality QA data sets can not only be used to fine-tune models, but also effectively evaluate the capabilities of large language models (LLM), especially the ability to understand and reason about scientific knowledge. Although there are currently many scientific QA data sets covering medicine, chemistry, biology and other fields, these data sets still have some shortcomings. First, the data form is relatively simple, most of which are multiple-choice questions. They are easy to evaluate, but limit the model's answer selection range and cannot fully test the model's ability to answer scientific questions. In contrast, open-ended Q&A

Machine learning is an important branch of artificial intelligence that gives computers the ability to learn from data and improve their capabilities without being explicitly programmed. Machine learning has a wide range of applications in various fields, from image recognition and natural language processing to recommendation systems and fraud detection, and it is changing the way we live. There are many different methods and theories in the field of machine learning, among which the five most influential methods are called the "Five Schools of Machine Learning". The five major schools are the symbolic school, the connectionist school, the evolutionary school, the Bayesian school and the analogy school. 1. Symbolism, also known as symbolism, emphasizes the use of symbols for logical reasoning and expression of knowledge. This school of thought believes that learning is a process of reverse deduction, through existing

Editor | KX In the field of drug research and development, accurately and effectively predicting the binding affinity of proteins and ligands is crucial for drug screening and optimization. However, current studies do not take into account the important role of molecular surface information in protein-ligand interactions. Based on this, researchers from Xiamen University proposed a novel multi-modal feature extraction (MFE) framework, which for the first time combines information on protein surface, 3D structure and sequence, and uses a cross-attention mechanism to compare different modalities. feature alignment. Experimental results demonstrate that this method achieves state-of-the-art performance in predicting protein-ligand binding affinities. Furthermore, ablation studies demonstrate the effectiveness and necessity of protein surface information and multimodal feature alignment within this framework. Related research begins with "S

According to news from this website on July 5, GlobalFoundries issued a press release on July 1 this year, announcing the acquisition of Tagore Technology’s power gallium nitride (GaN) technology and intellectual property portfolio, hoping to expand its market share in automobiles and the Internet of Things. and artificial intelligence data center application areas to explore higher efficiency and better performance. As technologies such as generative AI continue to develop in the digital world, gallium nitride (GaN) has become a key solution for sustainable and efficient power management, especially in data centers. This website quoted the official announcement that during this acquisition, Tagore Technology’s engineering team will join GLOBALFOUNDRIES to further develop gallium nitride technology. G
