For sustainability professionals, the task of calculating Product Carbon Footprints (PCFs) has long been the industry’s version of "accounting purgatory." It is a labor-intensive, data-hungry process that requires meticulously tracking the raw materials, energy consumption, and logistics of every individual SKU in a company’s portfolio.
However, a new wave of generative AI tools promises to transform this drudgery into a streamlined, automated workflow. By simply uploading a bill of materials or a product description, these platforms claim to generate instant emissions estimates. Yet, a new study from researchers at Watershed, a leading carbon accounting platform, serves as a sobering reminder: when it comes to climate data, speed should not be mistaken for accuracy. The study reveals that while AI can provide impressive "headline" numbers, the granular data beneath those figures is often riddled with significant, potentially misleading errors.
The Main Facts: The Promise of Instant Sustainability
At its core, a Product Carbon Footprint (PCF) is a measurement of the total greenhouse gas emissions generated throughout a product’s lifecycle—from the extraction of raw materials to manufacturing, distribution, and disposal. For corporations facing increasing pressure from regulators and consumers to decarbonize, these footprints are essential. They allow companies to identify "hotspots" in their supply chain where emissions can be reduced.
Despite their necessity, PCFs are notoriously difficult to scale. According to data from PwC, 69 percent of companies have managed to create PCFs for less than a quarter of their entire product lineups. The barrier is clear: collecting primary data from hundreds or thousands of tier-two and tier-three suppliers is a logistical nightmare.
Enter the AI-driven solution. Companies like Makersite and Terrascope are marketing tools that promise to bypass the tedious data-gathering phase. Makersite, for instance, boasts the ability to "automate accurate LCAs [Life Cycle Assessments] across your entire product portfolio in seconds," while Terrascope claims to reach 70 percent accuracy without needing direct supplier engagement. The pitch is enticing: in an era of rapid climate action, why wait months for supplier data when AI can provide an estimate in moments?
Chronology: The Evolution of Carbon Accounting
To understand why this shift is occurring, we must look at the timeline of corporate sustainability.
- Pre-2015: The Spreadsheet Era. Carbon accounting was a manual, error-prone process conducted primarily through fragmented Excel spreadsheets and static industry averages.
- 2015–2020: The Rise of Specialized Software. As ESG (Environmental, Social, and Governance) reporting became mandatory in various jurisdictions, specialized platforms emerged to help companies consolidate their data. However, the reliance on primary, vendor-specific data remained the "gold standard," keeping the process slow.
- 2023–2024: The Generative AI Boom. Following the public release of advanced Large Language Models (LLMs), the sustainability sector began testing whether these models could synthesize complex industrial data.
- 2026: The Watershed Study. A team led by Krishna Rao at Watershed conducted a landmark benchmarking test, pitting LLMs against expert-level carbon accounting benchmarks. This study marked the first time the industry began to systematically quantify the "hallucination" and logic gaps inherent in AI-driven PCF tools.
Supporting Data: When Accuracy Plummeted
The Watershed team designed a rigorous test to evaluate how well AI models—including top-tier engines from Anthropic, DeepSeek, Google, and OpenAI—could perform against established expert benchmarks. They tasked the models with estimating the PCFs for 175 products across a diverse array of sectors, including chemicals, textiles, and advanced electronics.
The results were, at first glance, promising. When looking at the "final answer," the best-performing models managed to land within two multiples of the expert-verified figure 77 percent of the time. For a high-level corporate slide deck, this might seem like a success.
However, the researchers dug deeper, looking at the intermediate steps of the AI’s logic. They asked the models to decompose the products into their constituent parts and estimate the emissions associated with each component. Here, the "intelligence" of the AI crumbled. The accuracy of these granular breakdowns plummeted to as low as 37 percent.
"What surprised me most was the size of the gap," Krishna Rao noted in the study. The implication is clear: the models were arriving at plausible final numbers through flawed, disjointed reasoning. In the world of accounting, this is a dangerous state of affairs. If the logic leading to the total is wrong, the model is essentially "right for the wrong reasons."
Official Perspectives and Industry Reactions
The industry reaction to these findings has been a mix of caution and calls for transparency. While the proponents of automated PCF tools argue that their software is more sophisticated than the generic LLMs tested by Watershed—often utilizing proprietary databases and "human-in-the-loop" oversight—the underlying message remains consistent: AI is a tool, not a replacement for expertise.

Companies like Makersite and Terrascope emphasize that their platforms are designed to provide "directionally correct" data, which is often sufficient for initial screening and hot-spot identification. They argue that the alternative—doing nothing because the process is too hard—is far worse for the planet than using an imperfect AI estimate.
However, researchers and sustainability auditors argue that "directionally correct" is a slippery slope. If a manufacturer relies on an AI estimate to justify switching vendors for a specific electronic component, and that estimate is based on flawed data, the company could inadvertently move to a supplier with a higher carbon footprint, effectively negating their sustainability efforts while claiming they have "optimized" their supply chain.
The Implications: A Call for "Explainable Sustainability"
The primary takeaway from the Watershed study is that sustainability professionals must move toward "explainable AI" (XAI) in their carbon accounting workflows. Users should never treat the final number as a black-box output.
1. The Risk of Misallocation
If a company allocates its capital or procurement budget based on AI-generated carbon data that contains systemic errors, they risk "greenwashing" their operations. If the intermediate steps of the AI’s calculation are hidden, the company cannot defend its data to auditors or stakeholders.
2. The Need for Human-in-the-Loop
The study reinforces the necessity of the "Human-in-the-Loop" (HITL) model. AI should be used to draft, categorize, and organize data, but the final validation must rest with human analysts who understand the nuances of industrial processes. An AI might assume a generic material weight, but a human engineer knows that a specific alloy used in a component significantly alters the energy intensity of the production process.
3. The Benchmarking Requirement
The methodology developed by Rao and his team at Watershed serves as a potential industry blueprint. Companies should not just ask for a software solution; they should demand to see the benchmark data. How does the provider’s model perform on specialized, verified datasets? If a vendor cannot demonstrate the accuracy of their intermediate calculations, the product should be treated as a draft, not a record of truth.
4. A Maturing Market
We are currently in a "hype cycle" for AI in ESG. As the market matures, we will likely see a bifurcation: platforms that offer "quick and dirty" estimates for general awareness, and platforms that offer "high-fidelity" accounting for regulatory reporting. The challenge for sustainability managers is to know which tool to reach for, and when.
Conclusion: The Path Forward
The automation of Product Carbon Footprints is an inevitable and welcome evolution in the fight against climate change. The sheer scale of the global supply chain makes manual accounting an impossibility for the vast majority of firms. However, as the Watershed study highlights, the current state of generative AI is not yet a substitute for the meticulous, rigorous work of expert carbon accounting.
For now, the mantra for sustainability teams should be "trust, but verify." By using AI to speed up the mundane parts of the process—such as categorizing components or identifying relevant emission factors—professionals can free up their time to focus on the high-level analysis that truly moves the needle. If we rely on AI to do the thinking for us, we risk turning our sustainability reports into works of fiction. If we use AI to support our thinking, we may finally have the tools we need to achieve the transparency and accountability that a warming planet demands.
Ultimately, the goal is not to have the fastest calculation, but the most accurate one. As the industry moves forward, the winners will be those who can harness the power of AI without losing sight of the technical, scientific, and logical rigor that true carbon accounting requires.



