Defining and assessing inventory record inaccuracy (IRI) metrics
Foreword
In retail, over 60% of inventory records in any one store have been found to be wrong, and when they are corrected and “trued up”, sales can grow by up to 11%. This is why we have prioritised research on the problem of inventory record inaccuracy metrics with a series of research projects that includes this new report that explores the different ways that retailers define and measure inventory record inaccuracy.
What the researchers have found is that in retail, there is no one standard measure or that any one retailer only uses one definition for Inventory Record Inaccuracy (IRI), and the set of guidelines they have produced should help retailers determine which metric should be used, for what purpose and goal. The interactive Excel spreadsheet accompanying this report has been designed to help retailers explore with their own data the different metrics.
Finally, in this research, the academics have shone a new light on the evergreen question, is the shrink problem the same as IRI? Their guidance is instructive and can help organisations articulate the similarities to, and the differences between IRI, defined in its many ways, and unknown loss, which is how retailers commonly, but not always, define shrink.
I would like to thank the Professors Glock, Rekik and Syntetos for carrying out this research and the retailers who helped contribute to the thinking and the findings presented in this report.
As with all the research undertaken on behalf of ECR Retail Loss, it would not be possible without the active support and involvement of the retail community and the many employees who generously gave their time to participate in the most meaningful of ways. Thank you for taking the time to share your thoughts and experiences – by working together we are much more likely to Sell More and Waste Less!
Finally, I encourage you to not only read and share this study but also to read the accompanying report that introduces a maturity and customer satisfaction model, that can act as a road map to organisations as they seek to tackle the IRI problem and unlock the lost sales potential. And our more recent report on predicting inventory record inaccuracy answers a valuable follow-up question: How can retailers move from merely identifying inventory record inaccuracy to predicting where it is most likely to occur and acting on it sooner?
John Fonteijn
Chair of the ECR Retail Loss Group
Executive Summary
Imagine that your organisation reports inventory accuracy of 95%. And another retailer claims accuracy of their records in the order of 90%. What does that tell you? Well, if the same metric has been used to report these results, then the answer is obvious: You do better than the competition. But if some different metric has been used, then any comparison is meaningless. Our research shows that retailers use a wide range of metrics to measure inventory accuracy, and that extreme caution is required when comparisons or benchmark exercises are undertaken not only across sectors and retailers, but also across categories and stores within an organisation.
In this report, we present the wide range of metrics that we found in use in the retail sector, and provide analysis of their advantages and drawbacks. We stress that in selecting an approach, simplicity should be prioritised over more complex and involved calculations, to ensure that the metrics are understood by everybody and can be easily applied. However, beyond this, there are also important logical and statistical factors that must be considered, in addition to practical managerial considerations, so that continuous and comparative data can be collected in order to inform decision-making. These are set out clearly in the report.
We find that retail practices towards the implementation of the metrics vary considerably, and although some retailers are applying the metrics correctly, in other cases application is associated with some serious logical shortcomings. Most notably, tolerances and percentage calculations are often made with regards to the system rather than physical stock, which is counter-intuitive: our point of reference should be what is actually happening (physical stock) as opposed to what we think is happening (system stock).
We also consider the interface between shrink and inventory record inaccuracy (IRI) metrics, noting that these terms should not be used interchangeably as they convey information about different issues. Ensuring consistency in the terminology used across different stakeholder groups is a major contribution of this report.
We believe that retailers need to appreciate the strengths and weakness of each metric in order to make informed decisions about when, where and how these metrics are applied. We are confident that this report goes some way towards this outcome. Thank you very much.
Introduction
Inventory record accuracy is of fundamental importance to the retail sector. Inventory records serve as the foundation for stock replenishment decisions, where inventory is typically reordered when recorded levels reach a predetermined replenishment threshold. Inventory records also play a critical role in omnichannel and e-commerce business models, enabling customers to access product availability information and reserve or order items online for in-store pickup or delivery. In these and other applications, it is essential that inventory records accurately reflect the actual stock available in the store or warehouse to ensure operational efficiency and customer satisfaction.
Prior research has shown though that more often than not, inventory records are wrong [1, 2]. Discrepancies can occur in both directions, i.e., the stock record can display more or less than is actually available in the store. Such inventory record inaccuracies (IRI) can stem from a variety of causes, including human errors, system limitations, theft, damage, and mismanagement of returns [1, 3]. Negative discrepancies, where recorded levels exceed actual stock, can result in delayed orders and stockouts, while positive discrepancies, where recorded levels fall short of actual stock, may lead to premature orders and overstocking [2]. Both cases have significant implications for operational efficiency, customer satisfaction, and financial performance [1, 2, 4]. In an era of heightened customer expectations, particularly with the rise of e-commerce and omnichannel retailing, consistent product availability is critical for retailers to remain competitive. Accurate inventory records underpin the ability to deliver on these expectations by ensuring that the right products are available at the right time and location.
The negative impact of IRI on business performance is also supported by prior research commissioned by ECR, in which we investigated the sales uplift associated with correcting inaccurate inventory records. That study, based on a field experiment comparing test stores subjected to a stock audit with control stores, found that sales could increase by 4% to 8% following the correction of inventory records [2].1 This finding highlights the commercial relevance of accurate stock data and the tangible benefits of addressing IRI.
Apart from impacting product availability, IRI may also affect store performance by introducing inefficiencies into operational processes. For example, inaccurate stock records can lead to wasted labour hours spent searching for non-existent items or manually reconciling discrepancies [2]. This diverts resources from value-adding activities, such as improving customer service or optimizing merchandising. Additionally, poor inventory record accuracy complicates financial reporting and decision-making, as inaccurate data can distort revenue projections, procurement plans, and cost analyses. For retailers operating on thin margins, these distortions can have serious consequences.
A follow-up study reported in [11] analyzed two stores of a single retailer and even found a sales increase of 11%.
To effectively address the IRI problem, retailers must rely on robust metrics to quantify the problem and guide their responses. IRI metrics provide critical insights into the extent of inaccuracies and help identify which product categories, Stock Keeping Units (SKUs), or locations are most affected. Without such metrics, retailers would lack the data necessary to prioritize interventions or evaluate the effectiveness of measures aimed at improving inventory record accuracy. By systematically monitoring IRI through well-defined metrics, retailers can track progress over time, assess the success of new processes or technologies, and ultimately align their efforts with broader operational and strategic goals.
While there is a growing body of literature on IRI that explores various dimensions of the problem, significant gaps remain in our understanding of how the retail sector measures IRI. Most of the existing literature and industry practice rely on simple binary accuracy measures – metrics that classify inventory records as either correct or incorrect, without regard for the size or direction of the error [5]. This framing also shaped our assumptions when we started looking at the issue of IRI measurement. However, and through our empirical investigation, it soon became evident that retailers often apply a much broader and (potentially) more sophisticated set of IRI metrics in practice. This includes not only binary measures, but also metrics that account for error size, direction, tolerance thresholds, and financial value – indicating that IRI measurement is an issue much more nuanced than commonly assumed.
Although several empirical studies incorporate metrics for measuring IRI (e.g., [6, 7, 8]), these metrics are often applied without substantial discussion or critical evaluation of their design or appropriateness. There are some comprehensive discussions of IRI measurement (see [4] and [9]) which also provide an overview of several metrics and examine key properties of effective IRI assessment. However, to date, there has been no research that systematically evaluates the metrics used by retail companies to quantify inaccuracies, or critically assesses alternative metrics towards enhancing industry practices. This lack of knowledge leaves both researchers and practitioners with no clear framework for understanding how IRI is measured, limiting the ability to benchmark performance or identify best practices. Addressing these gaps is critical to advancing the field and ultimately improving product availability at lower cost, and this is precisely what the project reported here is about.
We wish to understand how retailers perceive the assessment of inventory record inaccuracies, with a particular focus on the metrics they use to measure such inaccuracies and the practical ways in which these metrics are applied in daily operations. Through in-depth interviews with professionals around the world, we gathered qualitative insights into current measurement practices, their perceived value, and the challenges associated with implementing them. This provides a rich foundation for the analysis presented in this report.
The purpose of this report is twofold. First, it aims to provide a comprehensive overview of the metrics currently used in the retail sector to measure IRI. Second, it evaluates these metrics to identify their strengths, limitations, and practical applicability. By synthesizing these insights, the report seeks to equip retail professionals with practical guidance for selecting, applying, and improving IRI measurement as part of their inventory management practices.
Before we close this section, it is important to note that we make no attempts in this work to link IRI metrics to why or how IRI has happened in the first place. That is, we do not examine or infer anything about error root causes.
1 A follow-up study reported in [11] analyzed two stores of a single retailer and even found a sales increase of 11%.
Inventory Record Inaccuracy Measurement Practices
To gain insights into IRI measurement practices in the retail sector, we conducted semi-structured interviews with retail professionals. The research methodology we followed is explained in Appendix A. In the following section, each metric we identified is introduced with a plain-language explanation and a worked example to illustrate its application. This is followed by a brief evaluation of the metric’s practical relevance, including its advantages and potential limitations. Where relevant, we also highlight variations in how the metric is implemented or interpreted in different contexts, supported by examples and references where applicable. Formal mathematical definitions of the metrics are provided in Appendix B.
2.1. Signed Error
What it is:
This metric shows the total difference between the inventory that is physically in stock and what the system records – summed across all SKUs in a store, department, or product category. It builds on a basic SKU-level comparison (physical stock minus system stock), but the results are aggregated across product categories, departments or a store to provide a single overall value. The metric gives a high-level view of whether, on balance, inventory records tend to overstate or understate actual stock.
How it works:
To calculate the aggregate signed error, you:
- Add up the physical stock for all SKUs
- Add up the system stock for all SKUs
- Subtract the total system stock from the total physical stock:
Aggregate Signed Error = Total Physical Stock – Total System Stock
A positive result means that, overall, there’s more physical stock than expected; a negative result suggests there’s less. A total close to zero indicates that overstatements and understatements roughly cancel each other out (across the board).
Why it matters:
This measure gives a sense of the overall bias in your inventory records. A total close to zero suggests that overstatements and understatements in the system more or less cancel each other out. However, this can be misleading: even if every SKU is wrong, the net difference could still be zero. In other words, this metric might look “okay” even when individual SKUs are way off, so it’s not ideal for spotting specific problems, but it’s useful as a high-level indicator. The metric may be useful:
- as a quick health check for a store’s inventory records;
- when tracking shrinkage over time (especially if the result is negative);
- when reviewing stock accuracy from a financial perspective – for example, measuring whether the total value of inventory in the store matches what’s on the books.
Table 1. Example: signed difference calculated in units
| Item | Sales (in units) | System quantity | Physical quantity | Difference (in units) | Rate (Difference / Sales) |
|---|---|---|---|---|---|
| 1 | 40 | 20 | 19 | -1 | -2.5% |
| 2 | 200 | 45 | 40 | -5 | -2.5% |
| 3 | 400 | 65 | 65 | 0 | 0.0% |
| 4 | 200 | 42 | 38 | -4 | -2.0% |
| 5 | 800 | 90 | 95 | 5 | 0.6% |
| 6 | 600 | 100 | 91 | -9 | -1.5% |
| 7 | 150 | 45 | 42 | -3 | -2.0% |
| 8 | 350 | 78 | 15 | -63 | -18.0% |
| 9 | 1300 | 300 | 270 | -30 | -2.3% |
| 10 | 1700 | 815 | 920 | 105 | 6.2% |
| 11 | 100 | 30 | 0 | -30 | -30.0% |
| 12 | 160 | 0 | 45 | 45 | 28.1% |
| TOTAL | 6000 | 1630 | 1640 | 10 | 0.2% |
Notes: The total signed error across all twelve SKUs is +10 units, suggesting at first glance that system and physical inventory are closely aligned. However, this number masks large underlying discrepancies. For example, SKU 10 shows an overage of +105 units, while SKU 8 is understated by -63 units, and SKU 11 reports a full stockout with -30 units. These differences offset each other in the total, potentially creating a false sense of accuracy. This example also introduces a relative error rate, calculated as the signed difference divided by the number of units sold. While not part of the standard signed metric, this rate-based version is sometimes used in practice to contextualize discrepancies by SKU velocity or importance. However, as shown here, it can be highly sensitive to low sales volumes – resulting in extreme percentages for SKUs like 11 and 12. Please also note that the average result across SKUs is different depending on how it is calculated. If we are to take the total difference across SKUs divided by the total sales, this is 10 / 6000 = 0.2% (to the first decimal place). But if we are to average all rates across SKUs this would then be -2.2%. Overall, while the signed error can highlight directional trends (e.g., more understatements than overstatements), it remains ill-suited for assessing the true scale or operational impact of IRI.
Tips for interpretation:
This metric is potentially more meaningful when calculated in value terms rather than units, especially for accounting or shrinkage reporting. That way, large errors in high-value SKUs aren’t diluted by minor discrepancies in low-value items. Also, make sure your teams are consistent in how they calculate the underlying SKU-level differences – always confirm whether physical minus system, or the other way around, is being used.
2.2. Shrink
What it is:
This metric focuses on the loss of inventory that’s expected to be there (according to the system) but isn’t physically found. It helps quantify how much stock has gone missing relative to recent sales activity. It ignores any “gains” that have been discovered in the audit.
How it works:
Shrink is calculated only when the system shows more inventory than is actually in stock. If the system shows less, the error is ignored – this metric only tracks potential loss. At the SKU level, it is calculated as follows:
Shrink = (System Stock – Physical Stock) ÷ Sales × 100
(Only applies if System > Physical; otherwise, Shrink = 0)
The time interval over which sales are considered for inclusion in the shrink metric varies in practice, but typically this is the aggregated sales since the last stock audit. In terms of aggregating shrink, the following calculations are common. Please note that these calculations: i) lead to different results (see working example in Table 2); ii) apply when the system stock is greater than the physical one.
Average shrink per SKU:
Shrink (aggregate) = (Sum of shrink across SKUs) ÷ Number of SKUs, or
Shrink (aggregate) = Sum of (System Stock – Physical Stock) across SKUs ÷ Sum of Sales across SKUs x 100
Table 2. Example: shrink calculations
| Item | Sales (in units) | System quantity | Physical quantity | Negative difference (in units) = shrink | Shrink as a % (negative difference / sales) |
|---|---|---|---|---|---|
| 1 | 40 | 20 | 19 | -1 | -2.5% |
| 2 | 200 | 45 | 40 | -5 | -2.5% |
| 3 | 400 | 65 | 65 | 0 | 0.0% |
| 4 | 200 | 42 | 38 | -4 | -2.0% |
| 5 | 800 | 90 | 95 | 0 | 0.0% |
| 6 | 600 | 100 | 91 | -9 | -1.5% |
| 7 | 150 | 45 | 42 | -3 | -2.0% |
| 8 | 350 | 78 | 15 | -63 | -18.0% |
| 9 | 1300 | 300 | 270 | -30 | -2.3% |
| 10 | 1700 | 815 | 920 | 0 | 0.0% |
| 11 | 100 | 30 | 0 | -30 | -30.0% |
| 12 | 160 | 0 | 45 | 0 | 0.0% |
| TOTAL | 6000 | 1630 | 1640 | -145 | -2.4% |
Notes: If we take the ratio of the sum of the shrink across SKUs (here: -145 units) over the total sales (across SKUs; here: 6000 units), we get -2.4%. But if we are to average the shrink across SKUs (i.e., we sum up the individual shrink values appeared in the last column of the table and divide it by 12), we then get -5.1%.
Why it matters:
Shrink is a key concern in retail because it often reflects losses caused by theft, administrative errors, or unrecorded damage. This metric isolates cases where inventory has disappeared – that is, where the system shows more stock than is actually found. Expressing shrink as a percentage of sales puts the loss in
context, helping retailers understand how impactful discrepancies are relative to activity levels. This makes it especially valuable for:
- Loss prevention teams tracking stock loss by SKU or store;
- Shrink-focused reporting, especially after stocktakes or cycle counts;
- Store comparisons based on shrink performance and risk.
However, the term shrink is not always used consistently across the industry. In our interviews, we found that some companies use the term broadly, including not only losses (system > physical) but also overages (physical > system). While this approach may be intended to capture total discrepancy, it blurs the line between shrink and general inventory record inaccuracy, leading to potential confusion – especially when there is a positive number, implying a “gain” rather than a loss. Further, let us briefly say here that although shrink was reported as a metric in the interviews, it is atypical, in the sense that errors are reported as a percentage of sales (which is seldom the case for other relative metrics). It is because of these issues, that shrink warrants some further discussion and we return to this in Section 4 where we consider explicitly the interface between shrink and IRI.
Tips for interpretation:
We urge caution when treating shrink strictly as a loss-focused metric and reserving broader discrepancy calculations for the signed error metric, which accounts for both under- and overstatements. This avoids counterintuitive terminology and supports clearer communication across teams. For example, a store manager might see a “positive shrink” score and assume the store is over performing – while a loss prevention specialist might interpret the same result as cause for concern. Aligning definitions is especially important in cross-functional reporting and when comparing performance across regions or banners.
2.3. Binary error: percentage of inaccurate SKUs
What it is:
This metric shows the percentage of SKUs in a store, department, or category whose inventory records are incorrect. It’s a straightforward, high-level indicator of how widespread inventory record inaccuracies are, based on a simple yes/no test for each SKU.
How it works:
Each SKU is evaluated to check whether the system stock matches the physical stock exactly:
- If the physical and system stock match exactly, it’s counted as accurate;
- If there’s any difference, it’s counted as inaccurate.
- Then, the total number of inaccurate SKUs is divided by the total number of SKUs checked and multiplied by 100 to give a percentage:
% Inaccurate SKUs = (Number of SKUs with discrepancies ÷ Total SKUs) × 100
Table 3. Example: binary measurement
| Item | System quantity | Physical quantity | IRI |
|---|---|---|---|
| 1 | 20 | 19 | 1 |
| 2 | 45 | 40 | 1 |
| 3 | 65 | 65 | 0 |
| 4 | 42 | 38 | 1 |
| 5 | 90 | 95 | 1 |
| 6 | 100 | 91 | 1 |
| 7 | 45 | 42 | 1 |
| 8 | 78 | 15 | 1 |
| 9 | 300 | 270 | 1 |
| 10 | 815 | 920 | 1 |
| 11 | 30 | 0 | 1 |
| 12 | 0 | 45 | 1 |
| Stock record inaccuracy | 92% | ||
Notes: In this example, 11 out of 12 SKUs have wrong inventory records, resulting in a percentage of IRI of 11/12 = approx. 92%. The binary metric flags a record as “accurate” (IRI = 0 in the table) only if the system quantity and physical count match exactly. Even small discrepancies – such as a difference of just one unit – result in a score of one for that item. As a result, this metric gives a very strict view of inventory record accuracy. While the final percentage is easy to understand and useful for tracking general trends over time, it provides no information about the size or direction of the error. Here, item 3 is the only one considered accurate, even though several other discrepancies are minimal (e.g., item 1 differs by just one unit). This highlights a common limitation of the binary metric: it is highly sensitive to even small differences and may overstate the severity of stock record issues.
Why it matters:
This metric helps answer a simple but important question: “How many of our stock records are wrong?” Its strength lies in its clarity, ease of interpretation, and broad applicability across retail operations. Because the binary metric flags any mismatch – regardless of size – it is particularly useful for:
- Store and department benchmarking: helping identify variation and which areas maintain more or less accurate stock records;
- Trend tracking: enabling retailers to monitor whether data accuracy is improving or declining over time;
- Zero-error process enforcement: supporting systems where any discrepancy triggers immediate corrective action, making it ideal for real-time checks and automated workflows.2
The binary metric provides a straightforward, common language that can be shared across functions – from store staff to senior management – without requiring deep technical understanding. This makes it an excellent entry point for building awareness and accountability around inventory record accuracy. It also offers a straightforward way to compare performance across stores and in-between retailers.
Tips for interpretation:
- It doesn’t tell you how big the errors are;
- It doesn’t tell you if stock is over- or under-reported;
- It treats small and large errors the same way.
That said, the binary measure remains a practical and robust tool for high-level reporting, flagging where problems exist, and enforcing discipline in processes that demand precision. While more advanced metrics can offer greater insight into the nature of discrepancies, this measure plays a vital role in establishing a baseline and supporting simple, consistent decision-making. In addition, it has been proven an excellent measure in two further aspects. First, it can allow some benchmarking across retailers but also different industries / sectors. Distribution centres, for example, are known to report accuracy rates as % of SKUs being correct [10], and that can be contrasted to percentages reported in retail stores. Thus, the metric can serve as common language for retailers to compare between them but also to other sectors. Second, we know that the binary metric matters. Previous work we have conducted for ECR [2, 11] shows conclusively that truing up the inventory records (i.e. reducing the binary error rate) leads to sales increases between 4% and 11%. This is very important, as we do know that if we improve this metric, sales will grow.
2 In this context, automated workflows refer to system-driven processes that immediately trigger corrective actions – such as cycle counts, changes to the replenishment logic, or exception reports – whenever a stock record discrepancy is detected. This principle aligns with the Lean concept of Poka-Yoke (‘mistake-proofing’), which seeks to prevent errors or detect them at the earliest possible moment to avoid downstream process disruptions.
2.4. Error range: expressed as a percentage
What it is:
This metric is a variant of the binary error measure that checks whether the inventory error for a SKU is within an acceptable range, based on a percentage of what the system shows. If the error is small enough – within the defined tolerance – the record is considered accurate.
How it works:
Each SKU is assigned a tolerance level (e.g., 10%). The system then checks:
Is the error (in units) ≤ System Inventory × Tolerance %?
• If yes, the record is considered accurate;
• If no, it’s inaccurate.
Please note that tolerances may (and in our opinion should) be calculated on the physical rather than system inventories. However, the interviews tell us that retailers use more often the latter!
Table 4. Example: error range as percentage from system inventory
| Item | System quantity | Physical quantity | Absolute difference | 10% Error range | IRI |
|---|---|---|---|---|---|
| 1 | 20 | 19 | 1 | 2 | 0 |
| 2 | 45 | 40 | 5 | 5 | 0 |
| 3 | 65 | 65 | 0 | 7 | 0 |
| 4 | 42 | 38 | 4 | 4 | 0 |
| 5 | 90 | 95 | 5 | 9 | 0 |
| 6 | 100 | 91 | 9 | 10 | 0 |
| 7 | 45 | 42 | 3 | 5 | 0 |
| 8 | 78 | 15 | 63 | 8 | 1 |
| 9 | 300 | 270 | 30 | 30 | 0 |
| 10 | 815 | 920 | 105 | 82 | 1 |
| 11 | 30 | 0 | 30 | 3 | 1 |
| 12 | 0 | 45 | 45 | 0 | 1 |
| Stock record inaccuracy | 33% | ||||
1 Calculated with reference to the system quantity and rounded.
Notes: In this example, a tolerance of 10% was applied to each SKU. The stock record is considered accurate if the absolute error falls within that threshold, and inaccurate if it is outside. As a result, 8 out of 12 SKUs are classified as inaccurate, resulting in a percentage IRI of 33%. This stands in sharp contrast to the binary error metric, which considered only 1 SKU accurate under a strict “zero-difference” rule, yielding a 92% inaccuracy. The error range approach introduces practical flexibility. For instance, item 5 has a 5-unit difference on a system stock of 90 units – well within the allowed 9-unit margin. Instead of flagging this as inaccurate (as the binary metric does), the error range metric treats it as acceptable. This avoids overreacting to minor variances that may have little operational impact. However, this method still flags SKUs with significant discrepancies, like item 8 (63-unit difference) and item 10 (105-unit difference), as inaccurate – just as the binary metric would. Overall, the error range metric offers a balanced middle ground: more realistic than binary error for day-to-day reporting, but still sensitive enough to highlight where corrective action is needed.
Why it matters:
This is a more flexible alternative to the binary error metric. It recognizes that tiny errors – like 1 unit out of 100 – may not be operationally significant. It allows for a buffer, which makes reporting more meaningful and avoids overreacting to small discrepancies. It is especially useful
- For reporting SKU-level accuracy with tolerance for minor variances;
- In systems where accuracy thresholds differ by product type or category;
- To prioritize corrective actions – only flag what’s meaningfully “off”.
Tips for interpretation:
This measure is often based on the system inventory level, which may not reflect the actual stock – meaning it could incorrectly label some errors as acceptable. For greater accuracy, it would be better to define tolerance in relation to physical stock, not the system figure. It only makes sense to report an error in relation to what is actually happening (physical stock) as opposed to what we think is happening (system stock). We return to this issue later in the report, but for the time being please note that if we had used the physical rather than system stock to define the tolerances in the example in Table 4, the IRI would be different, 50%. Please also note that all comments regarding the measure’s performance apply regardless of whether the system or physical inventories are used to calculate the tolerances.
2.5. Error range: expressed in units
What it is:
This metric is a variant of the error range measure introduced earlier: it also checks whether the inventory error for a SKU is within an acceptable range, but it uses a fixed number of units as tolerance for each SKU instead of a percentage value. If the error falls inside this range, the stock record is considered accurate.
How it works:
Each SKU is given a tolerance level in units (e.g., ±2 units). The system checks:
Is the error (in units) ≤ Allowed Unit Tolerance?
- If yes, the record is marked as accurate;
- If no, it’s marked as inaccurate.
Table 5. Example: error range expressed in units
| Item | System quantity | Physical quantity | Absolute difference | Error range (in units) | IRI |
|---|---|---|---|---|---|
| 1 | 20 | 19 | 1 | 2 | 0 |
| 2 | 45 | 40 | 5 | 4 | 1 |
| 3 | 65 | 65 | 0 | 5 | 0 |
| 4 | 42 | 38 | 4 | 2 | 1 |
| 5 | 90 | 95 | 5 | 8 | 0 |
| 6 | 100 | 91 | 9 | 8 | 1 |
| 7 | 45 | 42 | 3 | 5 | 0 |
| 8 | 78 | 15 | 63 | 5 | 1 |
| 9 | 300 | 270 | 30 | 15 | 1 |
| 10 | 815 | 920 | 105 | 50 | 1 |
| 11 | 30 | 0 | 30 | 10 | 1 |
| 12 | 0 | 45 | 45 | 8 | 1 |
| Stock record inaccuracy | 67% | ||||
Notes: In this example, each SKU is evaluated using a fixed unit tolerance – for instance, a 2-unit threshold for item 1 or a 50-unit threshold for item 10. A stock record is considered accurate if the absolute error falls within the defined limit. The result: 8 out of 12 SKUs are marked as inaccurate in this example, yielding a stock record inaccuracy of approx. 67%. This unit-based tolerance approach is conceptually similar to the percentage-based version but operates with fixed thresholds rather than scaling them with SKU volume. This metric is useful when tolerances are based on operational impact or product criticality, but it requires careful tuning: a fixed threshold may underreport issues in low-volume SKUs and overreport them in high-volume ones. Compared to percentage-based tolerances, it offers more direct control but less scalability.
Why it matters:
This metric offers a straightforward way to handle small deviations without flagging them as issues. It’s especially helpful when you want to focus only on significant errors, and ignore the noise of small mismatches.
It is very useful
- In environments with low inventory levels (e.g., specialty or DIY retail);
- When most SKUs carry a similar volume of stock;
- When tolerances can be fine-tuned by product type or criticality (e.g., ABC classification).
Tips for interpretation:
Tolerance setting is tricky. Applying a blanket rule (e.g., “2 units for everything”) can be misleading, especially for high-volume SKUs where a 2-unit error is negligible. Unlike percentage-based tolerances, this measure doesn’t scale with SKU size – so it may over- or underreport inaccuracy depending on the item. When using this metric, consider combining this with ABC-type SKU classifications (e.g., A = high-volume or high-value, B = middle-volume or value, C = low-volume or value) and set tolerances accordingly. Alternatively, express tolerances in monetary value – i.e., flag records as inaccurate if the value difference between physical and system stock exceeds a set threshold. This may be more aligned with financial risk and loss prevention priorities.
2.6. Absolute error measure
What it is:
This metric shows the total number of units by which inventory records deviate from reality across all SKUs in a store, department, or category. It adds up the size of all individual errors, without letting positive and negative discrepancies cancel each other out.
How it works:
For each SKU, calculate the absolute difference between physical and system stock (i.e., ignore whether it is positive or negative, and just keep the number without sign). Then add up these values to reflect the total error across all SKUs. Because all differences are treated as positive, overstatements and understatements don’t cancel each other out – giving a clearer picture of the total inaccuracy:
Total Absolute Error = Sum of all |Physical – System| differences
Table 6: Example: absolute errors
| Item | System quantity | Physical counted | Absolute Difference |
|---|---|---|---|
| 1 | 20 | 19 | 1 |
| 2 | 45 | 40 | 5 |
| 3 | 65 | 65 | 0 |
| 4 | 42 | 38 | 4 |
| 5 | 90 | 95 | 5 |
| 6 | 100 | 91 | 9 |
| 7 | 45 | 42 | 3 |
| 8 | 78 | 15 | 63 |
| 9 | 300 | 270 | 30 |
| 10 | 815 | 920 | 105 |
| 11 | 30 | 0 | 30 |
| 12 | 0 | 45 | 45 |
| TOTAL | 1630 | 1640 | 300 |
Notes: The total absolute error across all ten SKUs is 300 units, reflecting the sum of all discrepancies regardless of direction. Unlike the signed error metric, this metric does not allow overages and shortages to cancel each other out – every deviation contributes to the total. For instance, item 10 has an absolute error of 105 units, and item 8 adds another 63 units, even though they had opposing signs in the signed version. This measure gives a clearer picture of the overall scale of inaccuracy in the system. While the signed error suggested only a small net difference (+10 units), the absolute error reveals substantial inventory mismatches. The metric is especially useful for understanding workload (e.g., reconciliation effort) or operational risk, though it still doesn’t distinguish between critical and non-critical errors. If retailers wanted to express this figure as a rate, they could divide the total absolute error (300 units) by the total physical stock counted (1,640 units), resulting in an absolute IRI rate of approximately 18%.
Why it matters:
This metric tells you how many units are wrong in total, making it much more meaningful than just summing signed errors, which can hide problems when over- and understatements balance each other out. This metric is particularly useful when the goal is to understand the scale of inventory problems rather than the direction. It helps answer questions like:
- How many units are our records off by, in total?
- Which stores or departments have the largest absolute discrepancies?
- Where should we focus audits or corrective actions?
It is often used in store-wide accuracy reviews, and operational planning – especially when large volumes of stock are involved and discrepancies can affect both customer experience and financial performance.
Tips for interpretation:
Unlike signed error metrics, this measure can’t “hide” problems through offsetting errors – every deviation is counted. That makes it a more reliable indicator of total error volume. Just remember: it treats all inaccuracies equally, regardless of operational impact. A one-unit error on a high-value item is weighted the same as a one-unit error on a low-value SKU. For this reason, many retailers calculate this metric in value terms as well
(e.g., euros or pounds), especially when setting performance targets or prioritizing loss prevention resources.
2.7. Mean Absolute Percentage Error (MAPE)
What it is:
The absolute inventory error for a specific SKU is expressed as a percentage of what the system says is in stock, to give the Absolute Percentage Error (APE), which gives a sense of how big the error is relative to expected stock. The Mean Absolute Percentage Error (MAPE) provides the average APE across a group of
SKUs, giving a single number that reflects how inaccurate your inventory records are overall, in relative terms.
How it works:
For each SKU, you calculate the Absolute Percentage Error (APE) – that is, how far off the inventory is, in absolute terms, as a percentage of what the system expected.
First, calculate the absolute difference between the physical and system stock. Then, divide that by the system stock, and multiply by 100 to get a percentage:
APE = (|Physical – System| ÷ System) × 100
(Note: This doesn’t work, i.e. the metric cannot be defined, if the system stock is zero, because we cannot divide by zero! Should this be the case, the relevant SKUs should be left out of the calculations.)
Then, you average the absolute percentage errors across all SKUs to obtain the MAPE:
MAPE = Sum of all APEs ÷ Number of SKUs
Table 7. Example: mean absolute percentage error
| Item | System quantity | Physical quantity | Absolute difference | Absolute percentage error |
|---|---|---|---|---|
| 1 | 20 | 19 | 1 | 5.0% |
| 2 | 45 | 40 | 5 | 11.1% |
| 3 | 65 | 65 | 0 | 0.0% |
| 4 | 42 | 38 | 4 | 9.5% |
| 5 | 90 | 95 | 5 | 5.6% |
| 6 | 100 | 91 | 9 | 9.0% |
| 7 | 45 | 42 | 3 | 6.7% |
| 8 | 78 | 15 | 63 | 80.8% |
| 9 | 300 | 270 | 30 | 10.0% |
| 10 | 815 | 920 | 105 | 12.9% |
| 11 | 30 | 0 | 30 | 100.0% |
| 12 | 0 | 45 | 45 | #DIV/0! |
| Mean Absolute Percentage Error (MAPE) | 22.8% | |||
Notes: SKU 12 is associated with zero stock in the system, and thus the APE cannot be defined. Excel will return an error if we attempt to get the ratio of the absolute difference over the system quantity. This would also be the case if we attempted to get the average APE (MAPE) across all SKUs inclusive of SKU 12. This is why SKU 12 was left out of the average calculation, to return a MAPE = 22.8%. If the physical quantity was to be used instead of the system one in the denominator, the APEs and MAPE would obviously be different.
Why it matters:
Expressing errors in relative terms can be powerful. At the individual SKU level, an absolute error of 5 units might be small for a fast-moving item with 500 units in stock (APE = 1%), but significant for an item where only 10 units are expected in the system (APE = 50%). APEs are very helpful in that way, and the MAPE is widely used to report inventory record inaccuracy across a store, department, or product category. It translates different unit errors into scale-independent percentages and helps prioritize problem areas – especially where small errors may be a big deal proportionally. It is especially useful for generating quick, understandable reports on overall inventory record accuracy, benchmarking performance across stores or teams, or identifying trends in record reliability over time.
Tips for interpretation:
Keep in mind that MAPE
- breaks down when the system inventory is zero, even for one SKU. In that case, the APE for that SKU cannot be defined and averaging across SKUs (say in Excel) will return an error. In cases like that, SKUs that are associated with zero system inventory should be left out of the calculations;
- suffers from an asymmetry problem: A 2-unit error looks worse if the system stock is smaller than the physical one. For example:
o If system stock = 4 and physical stock = 6 → APE = 2 / 4 = 50%
o If system stock = 6 and physical stock = 4 → APE = 2 / 6 = 33%
The unit error is the same, but the percentage isn’t. As such, APEs over-penalise positive IRI (i.e. physical stock being greater than the system one). Note that because of this asymmetry, some organizations prefer using a modified version of this metric – called the Symmetric APE (sAPE) – which we’ll explain next. The metric can be useful of course if the intention is to penalize positive IRI stronger than negative discrepancies;
- can be misleading as it bases the percentage reported on what the system says, not what’s actually in the store. This is a fundamental flaw, as logic dictates that discrepancies should be reported in relation to what is actually happening (physical inventory) as opposed to what we think is happening (system inventory). Some versions of this metric use physical stock in the denominator, or an average stock level over time. However, these are common in academic research but not often encountered (not by us at least) in retail operations;
- can be above 100%! This will happen every time when the physical stock is more than twice the system stock. For example, if the physical stock is (say) 25 units and the stock recorded in the system is 10 units, then the APE = 15/10 = 150%. And this is why it is dangerous to attempt to convert inaccuracy figures into accuracy. If the APE (inaccuracy) is (say) 20%, it is tempting to say that we have accuracy of 1 – 20% = 80%. But if the APE = 150% , then we cannot talk of accuracy of 1 – 150% = -50%.
2.8. Symmetric Mean Absolute percentage error (sMAPE)
What it is:
At the individual SKU level, this metric expresses the absolute inventory error as a percentage – but instead of dividing by the system stock (like APE does), it uses the average of the system and physical stock, leading to what is termed the symmetric APE (sAPE). Averaging the sAPEs then across a group of SKUs leads to the symmetric Mean Absolute Percentage Error (sMAPE). This metric was originally introduced (in the forecasting literature, see [12, 13]) to overcome the asymmetry problem of MAPE, though we will show that it suffers from a different asymmetry problem, perhaps more important than the one it was designed to tackle.
How it works:
For each SKU, you calculate the symmetric absolute percentage error (sAPE), that is, the absolute error is expressed in relation to the average of the system and physical inventory.
First, calculate the absolute difference between the system and physical inventory. Then, divide that by the average of the two values, and multiply by 100:
sAPE = (|Physical – System| ÷ ((Physical + System) ÷ 2)) × 100
Then, you average the sAPEs across all SKUs to obtain the sMAPE:
sMAPE = Sum of all sAPEs ÷ Number of SKUs
Table 8. Example: symmetric mean absolute percentage error
| Item | System quantity | Physical counted | Absolute Difference | Symmetric absolute percentage error |
|---|---|---|---|---|
| 1 | 20 | 19 | 1 | 5.1% |
| 2 | 45 | 40 | 5 | 11.8% |
| 3 | 65 | 65 | 0 | 0.0% |
| 4 | 42 | 38 | 4 | 10.0% |
| 5 | 90 | 95 | 5 | 5.4% |
| 6 | 100 | 91 | 9 | 9.4% |
| 7 | 45 | 42 | 3 | 6.9% |
| 8 | 78 | 15 | 63 | 135.5% |
| 9 | 300 | 270 | 30 | 10.5% |
| 10 | 815 | 920 | 105 | 12.1% |
| 11 | 30 | 0 | 30 | 200.0% |
| 12 | 0 | 45 | 45 | 200.0% |
| symmetric Mean Absolute Percentage Error (sMAPE) | 50.6% | |||
Notes: For both SKUs 11 and 12, the sAPE is 200%. Although this metric allows calculations to be performed when the system (or physical) inventory is zero, in such cases it will always return an error of 200%.
Why it matters:
At the individual SKU level, this metric overcomes the problem of division by zero but also the important asymmetry problem reported for the APEs:
- If system stock = 4 and physical stock = 6 → sAPE = 2 / 5 = 40%
- If system stock = 6 and physical stock = 4 → sAPE = 2 / 5 = 40%
The sMAPE is meant to give a “fairer” view of overall inventory error by avoiding the bias present in regular
MAPE – where overstatements and understatements are penalized differently. In theory, sMAPE could be helpful when you want to
• avoid the asymmetry problem of MAPE;
• include SKUs where the system stock is zero;
• get a more balanced summary of stock record inaccuracy across different products.
Tips for interpretation:
One issue with sMAPE is that it is not intuitive: Using the average of physical and system stock as a baseline for comparison feels artificial and doesn’t align with how most retailers think about inventory record discrepancies. In addition, and very importantly, it introduces a new type of asymmetry: it over-penalises negative IRI (i.e. physical stock being less than the system one). Consider the following example.
• System stock = 4, Physical stock = 6 → sAPE = 2 / 5 = 40%;
• System stock = 4, Physical stock = 2 → sAPE = 2 / 3 = 66%.
The absolute error in both cases is the same, but the metric suggests a higher penalty for the case where the physical stock is lower than the system one. Interestingly, in both those cases the APE would actually be 50%, unaffected by this issue:
• System stock = 4, Physical stock = 6 → APE = 2 / 5 = 50%;
• System stock = 4, Physical stock = 2 → APE = 2 / 5 = 50%.
Although sMAPE tries to improve on MAPE, it ends up being both harder to interpret and potentially more misleading. It does help when the system inventory is zero, as it leads to measurements that can be defined; however, in that case (and / or also when the physical inventory is zero) the sAPE will always equal 200%, and if there are many SKUs with zero (system/physical) inventory, then this would affect the sMAPE. There are further statistical arguments against the use of the sMAPE, the discussion of which is beyond the scope of this report. Due to these constraints and drawbacks, the sMAPE is not recommended for operational or reporting purposes.
2.9. Summary
In Table 9 we summarise the IRI resulting from the various metrics discussed in the report (Tables 1-8). It is obvious that different metrics mean different things and thus report completely different numbers. It is of paramount importance that retailers appreciate the plethora of metrics available to measure IRI and make informed decisions as to where and when they use such metrics.
The table goes also some way to show the scope for serious misunderstandings when percentages are used to compare performance across departments / stores / retailers, and sectors. Unless the same metric has been used to derive such numbers any comparison is meaningless.
Table 9. Summary of IRI reported under different metrics
| Metric | IRI |
|---|---|
| Signed Error | -0.20% |
| Shrink | -2.40% |
| Binary (Percentage) | 92% |
| Error Range (%) | 33% |
| Error Range (Units) | 67% |
| Absolute Error | 18% |
| MAPE | 22.8% |
| SMAPE | 50.6% |
As previously discussed, we further consider shrink and its interface with IRI in Section 4. Also, we have recently surveyed the literature and identified various metrics that are used in academic research, some of which are also reported here but some others don’t. For those interested in what academia has to say about IRI measurements, please refer to [5]. Finally, and to enable retailers to make the most of the material presented in this report, we complement the report with an electronic companion, an Excel file where all calculations are available and where colleagues may input their own data and find the respective IRI under the different measures discussed here.
Insights: using the metrics in practice
Inventory record inaccuracy is a multifaceted problem – and no single metric can capture all aspects of it. Different metrics highlight different things: some reveal the direction of discrepancies, others show the magnitude, while some simply count how many records are wrong. The key is to be aware of the strengths and limitations of each metric and be able to match them to your specific use case. This section outlines when each type of IRI measure is most useful, flags potential pitfalls, and offers guidance on combining metrics to get a fuller picture.
3.1. Replenishment optimization
When the goal is to improve replenishment decisions – either through automated systems or manual stock reviews – it’s essential to know not just whether records are wrong, but in which direction. Are items under- reported (risking stockouts), or over-reported (risking overstocks)?
Recommended metric: Signed error (in units or value).
Why: This metric captures the directional bias of the error. A consistently negative signed error (i.e., physical stock is lower than the system) may suggest hidden shrinkage or delayed updates – issues that can directly impair replenishment accuracy.
Caution: Signed errors can cancel out, masking big discrepancies. If SKU A is off by +50 and SKU B by –50, the total looks fine – even though both records are wrong. Use this metric alongside others if you need more detail, and pay attention to individual SKUs (which is what is needed anyway for replenishment purposes), particularly the high-volume and expensive ones.
Optional combination: Pair with a binary or tolerance-based metric to see how widespread the problem is, not just its direction.
3.2. Financial reporting and loss prevention
If you’re preparing stock valuations, loss reports, or shrink analyses, the focus shifts from accuracy per SKU to the total value impact. Here, knowing whether records are over or under matters less than understanding how much is missing, especially in monetary terms.
Recommended metrics: Shrink (% of sales or in value) or Aggregate absolute error (in units or value).
Why: These metrics quantify total loss or total deviation, which is exactly what’s needed for financial oversight. Shrink, in particular, focuses only on missing stock – aligning well with loss prevention priorities.
Caution: Be clear about definitions. In some organizations, “shrink” is used to describe both losses and overages – which can confuse stakeholders. We recommend using shrink strictly to indicate losses only and using signed error to describe net discrepancies. See also our discussion of shrink vs. IRI in Section 4.
Optional combination: Use absolute error (in value) to capture total deviation, and shrink to isolate losses. This gives a two-sided view: how much is missing, and how much is simply wrong.
3.3. Store performance benchmarking
Comparing stock accuracy across stores, departments, or regions requires a metric that’s simple, standardized, and easy to explain. For this, binary or range-based metrics are most effective.
Recommended metrics: Percentage of accurate SKUs (binary), Error range metrics (based on % or fixed unit tolerance), MAPE or sMAPE (if you want a relative view).
Why: These metrics provide a clean, comparable score – a percentage of records that pass or fail an accuracy threshold. This enables easy cross-store benchmarking, supports performance dashboards, and helps track improvement over time.
Caution: Binary metrics are strict – even a one-unit discrepancy marks a record as inaccurate. Range- based metrics are more flexible but depend heavily on tolerance settings. Setting too generous a range may hide issues; too strict may exaggerate them. This is why expressing (absolute) errors as percentages of the system (or physical inventory) as in MAPEs or sMAPEs may be preferable, though extreme caution is needed to ensure their asymmetries are properly understood.
Optional combination: Use a binary accuracy rate for reporting and MAPE or sMAPE to understand the average severity of errors across SKUs. This combination balances breadth and depth.
3.4. Operational root cause analysis
When you’re diagnosing why stock records are wrong – perhaps to improve receiving, gap scan processes, POS updates, or returns handling – you need more granular metrics that show which SKUs are affected, how badly, and in what direction.
Recommended metrics: Signed error per SKU, Absolute error per SKU, or Error distributions (histograms or counts by error band).
Why: These metrics let you zoom into the detail. A heatmap of signed errors, for example, may show that one department consistently undercounts, while another overstates. This helps pinpoint where operational breakdowns occur.
Caution: Detailed metrics need context. A 10-unit error may be major for one SKU and irrelevant for another. Always interpret numbers relative to volume, value, or sales.
Optional combination: Add tolerance-based IRI metrics to separate minor mismatches from systemic errors – so you don’t chase noise.
3.5. Cycle counting and store audits
In environments where ongoing counts are used to catch and fix errors, it’s crucial to flag records that matter most – e.g., high-value errors or those above a certain threshold.
Recommended metrics: Absolute error (in value or units) or tolerance-based IRI (especially fixed-unit range).
Why: These help prioritize which SKUs to count or investigate. For example, errors above €20 in value or more than 10 units can be flagged for review.
Caution: Thresholds must match operational goals. Too loose, and you miss important issues. Too tight, and you overwhelm staff with low-impact corrections.
Optional combination: Add error rankings (e.g., top 10 SKUs by error value) to direct attention efficiently.
3.6. Forecasting and demand planning
Inventory errors can distort demand signals. If you’re adjusting forecasts or safety stock settings, it’s important to detect persistent patterns of over- or understatement.
Recommended metrics: Signed error over time or shrink trends (monthly/quarterly).
Why: These show whether stock is being systematically misreported, which could bias demand models.
Caution: Requires time-series data. A one-off snapshot won’t reveal trends.
Table 10. Horses for Courses
| Use Case | Recommended Metrics |
|---|---|
| Replenishment optimization | Signed error (units or value) |
| Financial reporting | Shrink %, Aggregate Absolute Error (units or value) |
| Store performance benchmarking | Binary % Accurate SKUs, Range-based IRI, MAPE / sMAPE |
| Root cause analysis | Signed error, Absolute error, Error distributions |
| Cycle counting and auditing | Absolute error, tolerance-based IRI, Error rankings |
| Forecasting support | Signed error over time, Shrink trends |
Exploring shrink and IRI
In retail practice, the terms shrink and inventory record inaccuracy are sometimes used interchangeably, but they do not mean the same thing. This ambiguity is common across the industry and was highlighted repeatedly during our interviews and workshops: many retailers report shrink figures that include not only inventory losses but also situations where stock levels increased unexpectedly – for instance, when more items are delivered than ordered, or when goods are returned without being properly registered. While this practice is widespread, it creates conceptual confusion and can mislead stakeholders when interpreting results.
Another source of confusion is the widespread practice of expressing shrink as a percentage of sales. While this makes intuitive sense for loss prevention and financial reporting—putting missing stock in relation to business activity—it is problematic from an IRI perspective. Inventory errors are discrepancies between system and physical stock and are not inherently related to sales volumes. Two SKUs with the same unit level discrepancy can yield very different shrink rates if their sales differ, which can distort interpretations and comparisons. If financial impact is the goal, errors should instead be related to cost or value, not sales, as this creates a more meaningful link to decision-making.
At its core, shrink traditionally refers to a loss of inventory – stock that is expected to be present according to the system, but is missing when physically counted. This includes losses caused by theft, damage, spoilage, misplacement, or administrative errors that result in stock leaving the available assortment without proper booking. It is an important measure for loss prevention teams, store managers, and financial controllers, as it directly relates to lost sales opportunities and increased costs.
IRI, in contrast, is a broader concept. It describes any discrepancy between the inventory record and the actual stock, regardless of whether the system overstates or understates the quantity. From a measurement perspective, this includes both negative discrepancies (system > physical, indicating losses) and positive discrepancies (system < physical, indicating excess or unrecorded stock). Metrics such as the signed error capture both sides of this equation, showing the net difference between what is recorded and what is found.
The distinction matters because using shrink as an umbrella term for all discrepancies can lead to counterintuitive or even misleading reporting. For example, in the signed error dataset shown earlier in this report, we observed both significant losses (e.g., item 8: –63 units) and substantial overages (e.g., item 10:
+105 units). If both were summed across the entire dataset and reported as shrink, the net result (+10 units) would suggest a small gain – which does not reflect the underlying reality of large, offsetting discrepancies.
Loss prevention teams might interpret this as a positive outcome, while store managers could perceive it as a sign of good stock control, even though major operational and financial risks remain hidden. Our interviews revealed no consistent definition across the sector: some retailers restrict shrink reporting to losses only, while others include all discrepancies in their shrink figures. This lack of consensus can lead to misunderstanding in cross-functional settings. For instance, a merchandising team using shrink figures to plan replenishment might assume a positive shrink value indicates surplus stock, while finance or audit teams interpret the same number as an error to be corrected.
From both a conceptual and practical standpoint, we recommend adopting a clearer terminology framework:
• Use shrink strictly for losses (system > physical), as this aligns with its traditional meaning and common usage in loss prevention and financial reporting;
• Use signed error or inventory record inaccuracy (IRI) when referring to the net discrepancy including both losses and gains;
• For reporting that needs to capture the full scale of discrepancies (both directions), consider absolute error metrics or their value-based equivalents, which avoid cancellation effects while remaining neutral in terminology.
This clearer separation has several benefits: it supports better communication across teams, prevents false interpretations of “positive shrink,” and helps organizations design more targeted interventions. For example, shrink metrics can continue to drive theft prevention and damage reduction initiatives, while signed or absolute error metrics can inform replenishment planning, audit prioritization, and overall data quality improvement.
In summary, while shrink and IRI are closely related, they are not the same and should not be treated as interchangeable. Shrink is best understood as a subset of IRI – focusing specifically on losses – whereas IRI encompasses the full spectrum of inventory discrepancies. Aligning on these definitions helps ensure that the right metrics are used for the right purpose, and that organizational decisions are based on a shared understanding of what the numbers mean.
Conclusion
Companies are increasingly aware of the detrimental effects of inventory record inaccuracies on sales and also the way such inaccuracies lead to a range of serious inefficiencies, including non-value adding activities and resource waste. It is for these reasons that companies are increasingly investing heavily in stock record accuracy improvement. However, ‘improvement’ comes in many forms and it is extremely important that organisations appreciate the alternative ways available for measuring IRI, the advantages and disadvantages they each come with and their relationship with effective decision making.
Different companies use different metrics to measure IRI. This is not problematic per se, as long as companies:
i) appreciate that different metrics give different results which cannot then be compared, and ii) have made an informed selection of the metric they use, based on ease of use and interpretation, logical consistency, and link to decision making.
The purpose of this report is to equip retailers with the knowledge needed to achieve the above. To this end, we have the following key recommendations:
- Ensure that you are comparing like with like, meaning that if you wish to compare across categories and stores, the same measure must be used for/within each. Also be sure to do the same for external comparison and benchmarking purposes. In addition, be aware of the usually meaningless nature of absolute error comparisons, e.g. an error of 2 units for an SKU whose inventory averages at 200 units cannot be contrasted to an error of 2 units for another SKUs whose inventory averages at 10 units!;
- Ensure logical applications. For example, expressing an error as a percentage of the system inventory makes little sense. The aim is for errors to reflect reality and system inventory is what we think we have in stock, not what we actually do. Always express errors as a percentage of the physical inventory;
- Trial using simple numerical examples to “play” with the metrics and check for yourselves whether everything makes sense. Asymmetries (like those reported for the sMAPE) have detrimental effects for error measurement, and simple numerical examples are hugely valuable in exposing this serious shortcoming. The Excel file accompanying this report will be very useful as part of such an exercise, including allowing for the importing of your own data to explore how these play out using different metrics;
- Simplicity is important; however, the amount of information contained in the reported error and its link to decision making should also be considered when selecting an error metric.
These conclusions were reached through undertaking in-depth interviews with a large number of retail professionals with an explicit interest and involvement in IRI. Our investigation revealed not only that companies measure IRI in entirely different ways, but that their maturity towards such measurement exercises varies considerably. This report concerns the former of these, i.e. the metrics use. We consider the latter in a companion report also published by ECR.
We are confident that the findings and recommendations in this report will be of significant interest and value to the retail sector. Applying the recommendations set out above and in Section 3 will support companies in their journey towards greater inventory record accuracy. We would like to thank ECR for their support in the development and delivery of this important project and we look forward to continuing our work with the ECR community towards a more healthy and efficient retail sector.
Appendix A: Methodology
This appendix outlines the methodology used to obtain the results presented in this report. Given the limited research available on how inventory record inaccuracies are measured in the retail sector, this report draws on insights from a qualitative interview study conducted with retail professionals. The primary aim of the study was to explore how retailers perceive and address IRI and to develop a deeper understanding of the metrics and practices they use.
To capture a wide range of perspectives, the study employed semi-structured interviews. This approach allowed us to guide the conversation with predetermined, open-ended questions while also enabling flexibility to explore new topics introduced by the interviewees [14, 15]. Participants were recruited through multiple channels. First, a sequence of posts to those on the Efficient Consumer Response (ECR) Retail Loss mailing list invited retail professionals working in loss prevention, inventory control, or inventory auditing to participate in the study. Second, we directly contacted individuals from our professional networks. This dual recruitment strategy helped ensure a diverse pool of participants, reflecting variations in retail segments (e.g., grocery, fashion), business models (e.g., brick-and-mortar, online), geographical location, and seniority and job description of the interviewees (making sure in all cases that did possess the knowledge and information we were after).
In selecting interviewees, we followed the principles of ‘maximum variation’ and ‘replication logic’ [16, 17]. Maximum variation ensured that interviewees represented a broad range of organizational contexts likely to influence how IRI is measured and addressed. The replication logic allowed us to include similar cases to verify whether observations made at one retailer were consistent across others. The sample size was determined using the concept of ‘saturation’, where additional interviews were conducted until it became unlikely that new insights would emerge [18]. Saturation was reached after 25 interviews in our case, which included participants from a variety of roles and retail contexts. Table 11 provides an overview of the interviews conducted during the study.
Table 11: Overview of the interviewees participating in the study
| Role of the interviewee | Retail sector | Location |
|---|---|---|
| Retired Retail Expert | N/A | North America |
| Head of Global Store Operations | Grocery | Europe |
| Commercial Manager Loss Prevention | Household hardware | Oceania |
| Head of Stock Operations | Grocery | Europe |
| Performance Manager | Grocery | Europe |
| Product Director | Online Grocery | Europe |
| Supply Chain Developer | Fashion | Europe |
| Team Manager Stock Movement | Grocery | Europe |
| Internal Audit Manager | Grocery | Europe |
| Accuracy and loss business partner | Fashion | Europe |
| Project Manager Group Operations | Fashion | Europe |
| Team leader central supply function | Grocery | Europe |
| Head of Retail Operations | Pharmaceutical | Europe |
| RFID and Analytics Manager | Fashion | Europe |
| Retail Strategy Leader | Grocery | North America |
| Process Lead | Online Grocery | Europe |
| Senior Director Retail Operations | Grocery | North America |
| Stock Optimization and Retail Audit Manager | Household hardware | Europe |
| Solution Analyst Supply Chain | Grocery | Europe |
| Lead Analytics Manager | Grocery | Europe |
| National Director of Hypermarket Format | Grocery | Europe |
| Program Manager | Grocery | Oceania |
| Manager Loss and Fraud Prevention | Grocery | Europe |
| Head of Front Store Operations & Innovation Team | Pharmaceutical | North America |
| Replenishment Director | Grocery | Europe |
All interviews were conducted online using a video conferencing platform. Each session lasted between 20 and 70 minutes and was attended by two researchers from our team. With the consent of the interviewees, the interviews were audio- and video-recorded and subsequently transcribed using Sonix.ai. This process allowed us to capture detailed responses and ensure the accuracy of the data for subsequent analysis. Following best practices in qualitative research [18, 19], transcription and analysis commenced while the interview study was ongoing. This iterative approach allowed findings from earlier interviews to inform the questions and areas of focus in later sessions.
The transcripts were analysed using a qualitative coding process based on the Grounded Theory approach (see [20] for an example). This method involves developing theoretical insights directly (inductively) from the data rather than testing pre-existing hypotheses. Specifically, the analysis followed the conventional content analysis method described by [21]. Key thoughts were identified through repeated rounds of coding,and emerging categories were used to group related codes into clusters. Two researchers independently coded the transcripts and resolved any differences in interpretation through discussion, ensuring consistency and rigor in the analysis process. The semi-structured format of the interviews provided rich qualitative data, enabling us to explore a variety of perspectives and uncover unanticipated insights. By using open-ended questions and probing follow-ups [22], the interviews encouraged participants to provide detailed explanations and share practical examples of how IRI is managed and measured in their organizations. This approach enhanced the validity of the findings and ensured that the report reflects the complexities of IRI measurement and management in the retail sector.
- In addition to the interviews, we conducted a verification workshop in collaboration with ECR Retail Loss in March 2025. The session brought together 46 retail experts, many of whom had a background in loss prevention and some of whom had participated in the earlier interviews. During the workshop, preliminary findings from the interview study were presented and discussed in facilitated breakout groups. Participants provided feedback and contributed additional insights, which helped validate our findings and prompted minor refinements where necessary. This workshop served as a member validation exercise (cf. [23]), enhancing the credibility and practical relevance of our analysis through direct engagement with industry professionals. Following the inductive analysis and feedback session, we have worked on the preparation of this report and a complementary academic paper (cf. [5]) where we further elaborate on technical issues. Elaborating on the implications of our metrics analysis for retail practices is perceived as the main contribution of our work. This detailed methodology underpins the report’s objectives and findings, providing a robust foundation for understanding the metrics, practices, and contexts associated with IRI. Our methodology is graphically presented below.
Appendix B: Formal definition of IRI metrics
This appendix provides the technical details / mathematical formulations for the IRI metrics discussed in Section 2. We first introduce the notation used for formally representing the metrics and then provide the formal definition.
Notation
- t
- This stands for a specific time period — like a day, a week, or a month — depending on how often a retailer reviews inventory record accuracy.
- \( x_{k,t}^{IS} \)
- This is the number of units the system thinks are in stock for a specific product (SKU) k at the end of time period t.
We treat each SKU as unique to its store location — so a product in Store A and the same product in Store B count as two different SKUs.
- \( x_{k,t}^{PH} \)
- This is the actual number of units physically found in the store for the same SKU k at the end of time period t.
- \( q_{k,t} \)
- This is how many units of SKU k were sold during time period t.
- \( \alpha_k^{IS} \)
- This is a tolerance level, expressed as a percentage.
It shows how much of a difference between system and physical stock is still considered “acceptable” (for SKU k).
- \( \beta_k^{IS} \)
- This is another tolerance level, but in absolute terms — meaning how many units of difference are allowed between the system and physical stock before such discrepancy is considered a problem.
- K
- This is the total number of SKUs we’re looking at.
It could be all SKUs in a store or just a subset, like a product category or department.
Signed error
To calculate the signed error per SKU, subtract the system stock from the physical stock. In mathematical terms:
\[
IRI_{k,t}^{sign} = x_{k,t}^{PH} – x_{k,t}^{IS}
\]Signed error per SKUAggregate signed error across all SKUsTo calculate the aggregate signed error, you:
-
- Add up the physical stock for all SKUs
- Add up the system stock for all SKUs
- Subtract the total system stock from the total physical stock:
In mathematical terms:
\[
IRI_t^{sign\text{-}aggr}=\sum_{k=1}^{K}\left(x_{k,t}^{PH}-x_{k,t}^{IS}\right)=\sum_{k=1}^{K}x_{k,t}^{PH}-\sum_{k=1}^{K}x_{k,t}^{IS}
\]Aggregate signed error across all SKUsShrink is calculated only when the system shows more inventory than is actually in stock. If the system shows less, the error is ignored – this metric only tracks potential loss.
In mathematical terms:
\[
IRI_{k,t}^{shrink}
=
\frac{(x_{k,t}^{IS} – x_{k,t}^{PH})^{+}}{q_{k,t}}
\times 100
\]Shrink measure per SKU (positive inventory discrepancies only)Aggregate shrink measure across all SKUs (average across SKUs):
In terms of aggregating shrink, we simply average across SKUs, or more formally:
\[
IRI_t^{shrink\text{-}aggr}
=
\frac{1}{K}
\sum_{k=1}^{K}
\frac{(x_{k,t}^{IS} – x_{k,t}^{PH})^{+}}{q_{k,t}}
\times 100
\]Aggregate shrink measure across all SKUsBinary error
For each SKU, the system compares the physical stock to the system stock:
- If they match exactly, the error is 0
- If they don’t match, the error is 1
In mathematical terms:
\[
IRI_{k,t}^{bin}
=
\begin{cases}
0, & \text{if } \lvert x_{k,t}^{PH} – x_{k,t}^{IS} \rvert = 0 \\
1, & \text{otherwise}
\end{cases}
\]Binary indicator of inventory record inaccuracy per SKU% Inaccurate SKUs = (Number of SKUs with discrepancies ÷ Total SKUs) × 100
For each SKU:
-
- If the physical and system stock match exactly, it’s counted as accurate it’s counted as inaccurate
- If there’s any difference at all, it’s counted as inaccurate
Then, the total number of inaccurate SKUs is divided by the total number of SKUs checked and multiplied by 100 to give a percentage:
% Inaccurate SKUs = (Number of SKUs with discrepancies ÷ Total SKUs) × 100
In mathematical terms:
\[
IRI_t^{bin\text{-}aggr}
=
\frac{1}{K}
\sum_{k=1}^{K}
IRI_{k,t}^{bin}
\times 100
\]Aggregate binary inventory record inaccuracy across all SKUsError range: expressed as a percentage
Each SKU is assigned a tolerance level (e.g., 10%). The system then checks:
- Is the error (in units) ≤ System Inventory × Tolerance %?
- If yes, the record is considered accurate
- If no, it’s inaccurate
\[
IRI_{k,t}^{range,\%}
=
\begin{cases}
0, & \text{if } \lvert x_{k,t}^{PH} – x_{k,t}^{IS} \rvert \le \alpha_k^{IS} \cdot x_{k,t}^{IS} \\
1, & \text{otherwise}
\end{cases}
\]Range-based inventory record inaccuracy indicator (percentage tolerance)Then, the total number of inaccurate SKUs is divided by the total number of SKUs and multiplied by 100 to by the total number of SKUs and give a percentage:
% Inaccurate SKUs = (Number of inaccurate SKUs ÷ Total SKUs) × 100
In mathematical terms:
\[
IRI_t^{range,\%\text{-}aggr}
=
\frac{1}{K}
\sum_{k=1}^{K}
IRI_{k,t}^{range,\%}
\times 100
\]Aggregate range-based inventory record inaccuracy across all SKUsError range: expressed in units
Each SKU is given a tolerance level in units (e.g., ±2 units). The system checks:
Is the error (in units) ≤ Allowed Unit Tolerance?
- If yes, the record is marked as accurate 37
- If no, it’s marked as inaccurate
\[
IRI_{k,t}^{range,units}
=
\begin{cases}
0, & \text{if } \lvert x_{k,t}^{PH} – x_{k,t}^{IS} \rvert \le \beta_k^{IS} \\
1, & \text{otherwise}
\end{cases}
\]Range-based inventory record inaccuracy indicator using unit toleranceThen, the total number of inaccurate SKUs is divided by the total number of SKUs and multiplied by 100 to give a percentage:
% Inaccurate SKUs = (Number of inaccurate SKUs ÷ Total SKUs) × 100
In mathematical terms:
\[
IRI_t^{range,units\text{-}aggr}
=
\frac{1}{K}
\sum_{k=1}^{K}
IRI_{k,t}^{range,units}
\times 100
\]Aggregate range-based inventory record inaccuracy across all SKUs (unit tolerance)Absolute error measure
For each SKU, subtract the system stock from the physical stock, and take the absolute value
For each SKU, subtract the system stock from the physical stock, and take the absolute value (ignore whether it’s positive or negative):
\[
IRI_{k,t}^{abs}
=
\lvert x_{k,t}^{PH} – x_{k,t}^{IS} \rvert
\]Absolute inventory record inaccuracy per SKUFor each SKU, calculate the absolute error (as above), then sum those values across all SKUs. In mathematical terms
\[
IRI_t^{abs\text{-}aggr}
=
\sum_{k=1}^{K}
\lvert x_{k,t}^{PH} – x_{k,t}^{IS} \rvert
\]Aggregate absolute inventory record inaccuracy across all SKUsMean Absolute Percentage Error (MAPE)
First, calculate the absolute difference between the physical and system stock. Then, divide that by the Mean Absolute Percentage Error (MAPE) system stock, and multiply by 100 to get a percentage (APE)
\[
IRI_{k,t}^{APE}
=
\frac{\lvert x_{k,t}^{PH} – x_{k,t}^{IS} \rvert}{x_{k,t}^{IS}}
\times 100
\]Absolute percentage inventory record inaccuracy per SKU(Note: This only works when the system shows more than zero units.)
To calculate the MAPE then, obtain the average of the above across all pertinent SKUs:
\[
IRI_t^{MAPE}
=
\frac{1}{K}
\sum_{k=1}^{K}
IRI_{k,t}^{APE}
\]Mean absolute percentage inventory record inaccuracy across all SKUssymmetric Mean Absolute Percentage Error (sMAPE)
First, calculate the absolute difference between the physical and system stock. Then, divide that by the average of the system and physical stock, and multiply by 100 to get a percentage (sAPE):
\[
IRI_{k,t}^{SAPE}
=
\frac{\lvert x_{k,t}^{PH} – x_{k,t}^{IS} \rvert}
{x_{k,t}^{PH} + x_{k,t}^{IS}}
\times 100
\]Symmetric absolute percentage inventory record inaccuracy per SKUTo calculate the sMAPE then, obtain the average of the above across all pertinent SKUs:
\[
IRI_t^{sMAPE}
=
\frac{1}{K}
\sum_{k=1}^{K}
IRI_{k,t}^{SAPE}
\]Mean symmetric absolute percentage inventory record inaccuracy across all SKUsReferences
N. DeHoratius and A. Raman, “Inventory Record Inaccuracy: An Empirical Analysis,” Management Science, vol. 54, no. 4, pp. 627-641, 2008.
Y. Rekik, A. A. Syntetos and C. H. Glock, “Inventory Inaccuracy in Retailing: Does it Matter?,” ECR Retail Loss (available at: https://ecrloss.com/category/inventory-accuracy), 2019.
A. Shabani, G. Maroti, S. de Leeuw and W. Dullaert, “Inventory record inaccuracy and store-level performance,” International Journal of Production Economics, vol. 235, p. 108111, 2021.
C. Glock, A. Syntetos and Y. Rekik, “On the measurement of inventory record inaccuracies,” Working paper (available from the authors of this report), 2025.
S. Goyal, B. C. Hardgrave, J. A. Aloysius and N. DeHoratius, “The effectiveness of RFID in backroom and sales floor inventory management,” International Journal of Logistics Management, vol. 27, no. 3, pp. 795-815, 2016.
H. H.-C. Chuang and R. Oliva, “Inventory record inaccuracy: Causes and labor effects,” Journal of Operations Management, Vols. 39-40, pp. 63-78, 2015.
R. Ishfaq and U. Raja, “Empirical evaluation of IRI mitigation strategies in retail stores,” Journal of the Operational Research Society, vol. 71, no. 12, pp. 1972-1985, 2020. The Auburn University RFID Lab, “Inventory Accuracy,” available at: https://rfid.auburn.edu, 2023.
A. Marwane Kanoun, “Inventory Accuracy and Its Impact on Warehouse Performance. The case of Procter & Gamble.,” Politecnico di Torino, 2025.
Y. Rekik, R. Oliva, C. A. Syntetos and C. Glock, “Inventory record inaccuracy in grocery retailing: Impact of promotions and product perishability, and targeted effect of audits,” arXiv, (available from the authors of this report upon request) 2025.
J. E. Boylan and A. A. Syntetos, Intermittent Demand Forecasting: Context, Methos and Applications, Wiley, 2021.
P. Goodwin and R. Lawton, “On the asymmetry of the symmetric MAPE,” International Journal of Forecasting, vol. 15, pp. 405-408, 1999.
E. Grosse, S. M. Dixon, W. P. Neumann and C. H. Glock, “Using qualitative interviewing to examine human factors in warehouse order picking: technical note,” International Journal of Logistics Systems and Management, vol. 23, no. 4, pp. 499-518, 2016.
R. K. Yin, Case Study Research and Applications: Design and Methods, 6 ed., Los Angeles: Sage Publications, 2018.
N. Anand, H. K. Gardner and T. Morris, “Knowledge-based Innovation: Emergence and Embedding of New Practice Areas in Management Consulting Firms,” Academy of Management Journal, vol. 50, no. 2, pp. 406-428, 2007.
K. M. Eisenhardt and M. E. Graebner, “Theory Building from Cases: Opportunities and Challenges,” Academy of Management Journal, vol. 50, no. 1, pp. 25-32, 2007.
K. M. Eisenhardt, “Building Theories from Case Study Research,” Academy of Management Review, vol. 14, no. 4, pp. 532-550, 1989.
J. M. Corbin and A. Strauss, “Grounded Theory Research: Procedures, Canons, and Evaluation Criteria,” Qualitative Sociology, vol. 13, no. 1, pp. 3-21, 1990.
R. N. Raghuraman, A. Upasani, A. Gonzales, J. Aviles, J. Cha and D. Srinivasan, “Manufacturing Industry Stakeholder Perspectives on Occupational Exoskeletons: Changes after a Brief Exposure to Exoskeletons,” IISE Transactions on Occupational Ergonomics and Human Factors, vol. 11, no. 3-4, pp.
H.-F. Hsieh and S. E. Shannon, “Three Approaches to Qualitative Content Analysis,” Qualitative Health Research, vol. 15, no. 9, pp. 1277-1288, 2005.
K. Roulston, “Considering Quality in Qualitative Interviewing,” Qualitative Research, vol. 10, no. 2, pp. 199-228, 2010.
K. Thoring, R. Mueller and P. Badke-Schaub, “Workshops as a research method: Guidelines for designing and evaluating artifacts through workshops,” Proceedings of the Annual Hawaii International Conference on System Sciences, pp. 5036-5045, 2020.
Disclaimer
The research for this report was supported by the ECR Retail Loss Group. The report is intended for general information only; it is based on a review of available literature together with primary research undertaken with retail organisations worldwide. Individuals or companies are advised to seek professional guidance regarding their specific needs and requirements before taking any actions based on anything contained in this report. Any such actions taken by individuals or companies are entirely at their own risk. Companies are also responsible for ensuring they comply with all relevant laws and regulations, including those relating to intellectual property rights, data protection, and competition laws or regulations. The images used in this document do not necessarily reflect the companies taking part in this research. © November 2025, all rights reserved.
About the Authors
Christoph H. Glock is head of the Institute of Production and Supply Chain Management at Technical University of Darmstadt, Germany. His research focuses on the coordination of inventory replenishments and the management of physical stocks in warehouses. He has worked with many companies, and decision support models and methodologies co-developed by him are successfully used in industry to manage inventories and warehousing operations. He is a member of several professional societies and an editor of three international scientific journals.
Aris A. Syntetos is Distinguished Research Professor of Decision Science and the DSV Chair of Logistics at Cardiff Business School, Cardiff University. He researches forecasting and uncertainty management to improve inventory and supply chain decisions. He has worked with numerous organisations worldwide, and many of his inventory-forecasting algorithms are used in major supply chain software packages and in-house solutions, delivering significant economic benefits. He is past Director of the International Institute of Forecasters and current Vice President of the International Society for Inventories Research.
Yacine Rekik is Professor of Operations and Supply Chain Management at emlyon business school. His research focuses on the digital transformation of supply chains, including inventory management, record inaccuracy, RFID-enabled systems, omnichannel operations, and AI-driven decision models. He holds a PhD from École Centrale Paris and an HDR from INSA Lyon, was a Research Associate at Cambridge’s DIAL, and collaborates with the ECR Retail Loss Group.
Acknowledgements: The authors thank all companies that participated by sharing their experience and knowledge in interviews and workshops. They also thank retailers who provided feedback during meetings of the ECR Retail Loss Group. Special thanks go to Stephan Kolassa from SAP for his constructive comments that helped improve an earlier version of the report.
To contact the authors: glock@pscm.tu-darmstadt.de; aris@cardiff.ac.uk; yrekik@em-lyon.com
The same academic team also runs inventory record accuracy training where the three Professors look at practical implications of their research. The 2026 IRI Training course in Cardiff on 9 and 10 December covers the drivers of stock record inaccuracies, how to articulate the business case for improvement, countermeasures to enhance accuracy, and how to use research insights in practice. It is designed for loss prevention, supply chain and operations leaders and uses interactive exercises and gamification alongside academic instruction.
About ECR Retail Loss: ECR Retail Loss is part of ECR Community, a voluntary retailer-manufacturer platform focused on fulfilling consumer wishes better, faster, and at less cost. For over 20 years, the Group has acted as an independent think tank developing new ways to manage loss and on-shelf availability in the retail industry. Championing the idea of “Sell More and Lose Less,” the Group is open to any retailer or manufacturer. For further information: ecrloss.com
Research commissioned by the ECR Retail Loss Group is made possible by independent research grants provided by Axon, Checkpoint Systems, NCR Voiyix, Retail Insight, RGIS and Vusion.

















