20 September 2025

The Land of Startups is a Land of Pivots

We've graduated from the Peachscore+Dealum Accelerator, and in the midst of another pivot; having identified a beachhead industry for our data operations solution. The blog will be on hold for this quarter. 

19 August 2025

Converting intangibles to operating assets – Carbon Credits vs. Data


By David Huer | Aug 2025

Carbon Credits as a Universal Commodity: From Standards to Specifications, E*Comdty Research & Advisory, Capitalism, Freedom, and Carbon Credits, March 21, 2025 [Link]

Valuation Approaches:

The article argues that developers of process-driven innovations should not seek external standards verification--nor be compliant to those standards "in an effort to measure legitimacy". Instead, the author calls for adherence to "metrics (that are) tied to tangible, provable processes" that deliver repeatable, verifiable outputs.

Case:

The core challenge with carbon credits is their economic structure. A "carbon credit" is meant to function as an asset, yet its definition is blurry, its property rights weak, and its market adoption uneven. This makes it difficult for firms to treat credits as true operating assets, which in turn slows investment and undermines trust. 

The article argues that carbon credits valuation should evolve to process-based proof rather than external, shifting standards. Verification should rely on measurable, quantifiable outcomes (e.g., CO₂ captured, plastic recycled), transparently recorded through blockchain and digital ledgers. This enables tokenization—turning verified credits into standardized, fungible tokens for efficient trading across exchanges and OTC markets. Specifications emerge organically from market consensus, ensuring liquidity and transparency, unlike imposed standards that create inefficiencies. 

DreamWorks’ securitization of copyright license receivables in the intangible asset space is mirrored conceptually here: carbon credits become credible when process-based, digitized, and traded under market-defined specifications. A practical example is blockchain-enabled carbon removal verification, where X tons of CO₂ removed are transparently linked to Y tradable credits.

Implications for Financial and Strategic Reporting:

The suggested approach is that companies can adopt process-based verification to generate measurable, auditable outcomes. Tokenized carbon credits can be treated as standardized assets, enhancing balance sheet clarity, risk reporting, and capital allocation. For regulators, the role becomes that of referee—enforcing transparency and fraud prevention, but not dictating methodologies. This reframing aligns carbon credits with other commodities, integrating them into corporate reporting frameworks and investor-grade ESG disclosures.

Key Takeaways:

  • Specifications driven by process-produced evidence enable liquidity and transparency.
  • Blockchain and tokenization embed proof directly into carbon credits, making them self-verifying and universally tradable.
  • Governments should enforce legal and fraud protections but avoid dictating methodologies.
  • The transition to tokenized credits faces scaling challenges but offers an elegant solution for creating a universal, trusted carbon credit market.
  • Ultimately, proof of process—not compliance to standards defined to refuse true innovation—will drive legitimacy and market adoption.

09 July 2025

Valuing Corporate Intangibles








By David Huer | June 2025

(2023) Nicolas Crouzet, Yueran Ma, FINANCING AND VALUATION OF INTANGIBLE ASSETS, Nicolas Crouzet, Yueran Ma, Kellogg School of Management, Northwestern University, [Link] (Accessed Q2-2025). Expert Consultative Group on Valuation of Intangible Assets, GENEVA, OCTOBER 12, 2023

Summary

The paper discusses methods to value intangible assets; and opens by emphasizing the growing utility of intangible assets modern companies. These assets are increasingly vital but present major challenges for financing and valuation. 

Crouzet and Ma note the key distinction between separable intangibles (e.g., patents, software, brands) that can be sold or pledged independently, and nonseparable intangibles (e.g., know-how) that are deeply embedded in the firm and which is  typically financed through equity or enterprise-level debt. 

Structural bottlenecks limit financing use cases options for intangibles: unclear property rights, thin markets, and limited accounting transparency. For separable assets, valuation is hindered by the scarcity of secondary market transactions. The context of  the valuation is also to be considered: 

(1) Economists: "...define intangible assets as non-physical, firm-controlled resources resulting from past expenditures that are expected to yield future economic benefits ( proprietary knowledge, software, customer relationships, and internal databases. Crucially, this definition excludes financial assets and public goods like open-source software, as these do not stem from exclusive firm investment."

(2) Financial Accountants: "...require identifiability: assets must be separable or arise from contractual/legal rights to be recognized. As a result, only acquired intangibles are typically recorded on balance sheets, while internally developed assets are often omitted." 

The authors continue by delving into the role of uncertainty and discount rates in valuation.

The authors argue that in this age, it is prudent to capture "the true scope and value of firm-created intangible capital." One of the ways is to ensure the presence of "Robust institutional support—especially bankruptcy protections and accurate reporting—(which are) is essential for unlocking capital derived from intangible assets.  

This is challenging, so  income-derived valuation methods are often used. Methods are used on a case-by-case basis. Moreover, the authors "caution against inflating discount rates based on perceived risk; instead, systematic risks should influence discounting, while idiosyncratic risks, such as obsolescence or legal uncertainty, are better handled by adjusting cash flow projections. Income-derived methods include: With-and-without, Relief-from-royalty, Excess earnings, and Greenfield method. These are explained in the article.




06 June 2025

China’s Treatment of Data Assets






By David Huer | June 2025

Recognizing Data as Inventory: A Step Toward Accounting Clarity

(2022) Xiong, F., Xie, M., Zhao, L., Li, C., & Fan, X. (2022). Recognition and Evaluation of Dats as Intangible Assets. SAGE Open, 12(2): https://doi.org/10.1177/21582440221094600

A Longstanding Gap in Financial Reporting:

As covered previously in this blog, internally generated data is routinely used to drive advertising, pricing, and business strategy—yet is often absent from balance sheets. Under IFRS and GAAP intangible assets must be identifiable, controlled, and expected to generate future economic benefits. Despite many datasets meeting these criteria, recognition has lagged due to challenges in valuation and disclosure practices.

In this 2022 article, Feng Xiong and colleagues proposed that enterprise data should be explicitly recognized under China’s accounting standards—similar to the treatment already permitted by IFRS and GAAP. Their analysis provides a structured approach to valuation and calls for formal inclusion of qualifying data assets in financial reporting.

Valuation Approaches:

The authors reviewed three standard methods for valuing data assets:

  • Cost Approach: Captures direct and indirect costs to collect, clean, and maintain data. While straightforward, it may undervalue reusable or high-impact datasets.

  • Net Present Value (NPV): Estimates the discounted future benefits of data-driven activities. This is particularly relevant for companies with subscription models or recurring data-based services.

  • Market Approach: Derives value from data marketplace transactions or similar third-party pricing. This method is still developing, especially for proprietary or non-tradable datasets.

Case Example: Hithink RoyalFlush

Hithink RoyalFlush, a Chinese fintech firm, generates substantial revenue from data-driven services, including telecom tools, advertising, and software subscriptions. Yet its financial statements did not list data assets. Given the proprietary nature of its data, the market approach is unsuitable. The cost method underrepresents value due to repeated reuse. The authors proposed applying the NPV-based approach to more accurately reflect the strategic role and long-term value of these assets.

Implications for Financial and Strategic Reporting:

Recognizing internally generated data as intangible assets can improve reporting accuracy and better reflect a company’s economic reality. Benefits include:

  • Closer alignment between business value and reported assets
  • Improved transparency for investors and analysts
  • More robust valuation in mergers and acquisitions
  • Greater support for monetization and licensing strategies

However, recognition must be grounded in reliable valuation and meet existing accounting criteria, especially regarding measurability and probability of future benefit.

Key Takeaways

The authors argue that recognizing data as intangible assets improves the accuracy and usefulness of financial reporting, particularly for firms whose business models rely heavily on data analytics. They emphasize that existing accounting principles already support this recognition, and propose three valuation methods—cost, NPV, and market-based—each suited to different business contexts. 

They call for further research to refine when and how such recognition should apply, and they caution that internal auditors will play a critical role in ensuring firms apply these principles ethically and within legal boundaries. Ultimately, their work advocates for accounting systems that reflect the strategic importance and monetizable value of data in the modern digital economy.

------------------

Postscript: Regulatory Developments

In 2023, China implemented new accounting rules to allow enterprise data to be included on the balance sheet (1). The change is designed to help companies more accurately reflect the financial value of their data assets—particularly in the context of sales or licensing.

The new approach defines data as an intangible asset that is organized into two categories:

  • “Intangible Assets”: Data used internally for strategic or operational purposes

  • “Inventory”: Data packaged for sale (e.g., APIs, datasets, analytic products): Note: This classification is an accounting mechanism to support reporting and valuation. It allows companies to recognize saleable data assets as “inventory” for the purposes of revenue tracking, cost-of-goods-sold (COGS), and gross margin accounting. 

  • Data's fundamental nature ("intangibility") does not change. 

(1) https://www.forbes.com/councils/forbestechcouncil/2024/04/18/china-treats-data-as-an-asset-heres-why-your-business-should-too/

05 April 2025

2025 Q2 Pivot Note Update


The future looks bright.


Hello everyone, our pivot continues; and now forecast re-starting blog posts in May-June. This will be a review of a Chinese data valuation paper: https://journals.sagepub.com/doi/full/10.1177/21582440221094600

These two "Intangible assets" papers are also in pipeline: 


Accelerator Recommendation:

  • The Peachscore+Dealum Accelerator (NYC/Silicon Valley) has been an exceptionally good experience. 
  • If you are a venture founder, take a look and consider an application: https://peachscore.com/ 



13 January 2025

2025 Q1 Pivot Note

 Hello everyone, we are in the midst of a pivot; and forecast resuming blog posts at the start of March or April 2025. In the meantime, here's our corporate card for 2025. Best of the year to everyone.



01 December 2024

Towards Standardization of Data Licenses: The Montreal Data License, Benjamin et al (2019)

Overview: The brief accomplishes two tasks. 1st) It explores the intellectual underpinnings that prevent full use of data by: (i) market participants who want access to data without paying the creators of that generated data, and: (ii) for creators who want to make their data freely available without loss of control of data rights to others. 2nd) The authors go on to propose a novel data licensing regime to address the shortcomings of current approaches.


The regime "focuses on contracts for accessing databases rather than recognizing specific legal statuses for databases." It is a good first step towards creating a rigorous regime for open source data usage licenses.





  • Part (1) Introduces policy interest in unambiguous data licensing
  • Part (2) Describes current licensing regimes
  • Part (3) Describes taxonomy regime to activate use of MDL
  • Part (4) Discusses caveats and exceptions
  • Part (5, 6, 7) Presents Conclusions, References & Appendices).


Comment: The Montreal Data License (MDL) is a first iteration of an approach that tackles the problem of licensing data as a free good. The authors propose to standardize terminology and licensing standards, to make data available in a manner that is similar to the Free & Open Source Software licensing regime. The authors note that “While metadata can help reduce some of these costs, it often lacks the legal clarity that the needed to define how data can be used.” In this context, the authors propose MDL to create economic conditions that “enable AI and machine learning (ML) growth that benefits everyone.”


There are two items of concern:


There is a claim that “progress made in ML and AI should also be reflected by licenses that reflect the iterative process that move fundamental research to commercially available products and solutions, akin to the foundational moments of license creation for FOSS;” and “For the benefits of ML and AI to be accessible and benefit a wider realm of humanity, other market participants need to be on a level playing field.” 

This affects the value of the paperIt is reasonable to create “a more transparent, predictable market for data with clear legal language as its underpinning” But there does not seem to be the consideration of the rippling effects of the proposed regime. For example: 

    • Issue #1 - Market Effects: This proposal appears to support the idea that markets ought to be manipulated to create a falsely-constructed "level playing field" that is designed to favour a single class of market participant (fundamental researchers). If this is what is being suggested, that might not be acceptable to other market participants.
    • Issue #2 - Ecological/Emissions Effects of False Equivalency: FOSS data goods are not equivalent to FOSS software goods. FOSS data goods are more likely to have no owner, or lose their owner, and thereby contribute to the blight on the planet that is waste data [Note: OrbMB is developing a process to help cut data waste holdings - Ed.]. 


It would be prudent to develop supplementary licensing frameworks to address these issues. For example: The MDL FOSS-Data Regime interest group might want to consider the nature of custodianship--does a FOSS Data licensee accept legal responsibility for chain-of-custody, cost management, emissions cost management, and duty-to-delete the asset copy? 

===============================

(2019) Misha Benjamin et al, TOWARDS STANDARDIZATION OF DATA LICENSES: THE MONTREAL DATA LICENSE, Misha Benjamin1, Paul Gagnon 1; Negar Rostamzadeh 1; Chris Pal 1,2,3; Yoshua Bengio3,4,5; Alex Shee 1. 1 Element AI; 2 Polytechnique Montréal; 3 MILA; 4 Canada CIFAR AI Chair; 5 Senior CIFAR Fellow, arXiv:1903.12262v1 & https://arxiv.org/abs/1903.12262  (Accessed Q3/4-2024)

===============================

Part (1) Introduction (p.1-2)

This paper introduces a taxonomy for data licensing in artificial intelligence (AI) and machine learning (ML). The aim is to create a standardized framework similar to open-source software licensing.

Drawing a parallel between data and oil markets, the paper highlights the critical role of data in powering AI and ML systems; comparing the absence of standardized and regulated data acquisition and processing to the resource-intensive nature of oil extraction, refining and delivery. Unlike O&G markets, the data market lacks such frameworks, which the authors propose require heavy regulation to reduce friction, ensure security, and build public trust. 

The authors argue that a new licensing regime will create “fairer and more efficient markets for data” as the new regime will more clearly “define how data can be used in the fields of AI and ML.” A “new family of data rights is organized as a new form of license called the Montreal Data License (MDL) and there is a web-based tool for generating these licenses.

While metadata can help reduce some of these costs, it often lacks legal clarity to properly define how data can be used. The authors call for broader access to AI benefits, claiming that a more transparent, predictable data market with clear legal language is needed. This paper discusses the challenges of current data licensing terms and proposes a taxonomy that better aligns with AI and ML, aiming to clarify the use of data and associated rights, providing a framework for database creators to generate clearer licensing terms; to provide access to data for research purposes.

===============================

Part (2) Licensing barriers to use of data in ML and AI (p.3-5)

A review of commonly used databases in AI research (Appendix 1, Document page 12) reveals a "patchwork" of vague licensing terms that create uncertainty about the permissions granted for their use. This gives rise to barriers to usability such as:

(i) Lack of Nuance on “Use”: The right to use is granted, without defining what “use” actually means; this creates “one homogenous notion of use” and this creates downstream problems.

(ii) Commercial vs Non-Commercial Use: By way of example, many of the licenses cited in Appendix 1 contain a restriction against commercial use. It is the opinion of the authors that this is problematically ambiguous;

(iii) Barriers to Research: Pure-play academic researchers are increasingly unable to access datasets because of ambiguity and cost.

(iv) Lack of Uniformity: Terminology is not uniform and standardized. Free and Open Source Software (FOSS) communities have built a software-sharing commons by constructing a standard terminology and rules for software use. Data might move to the same regime. 

(v) Share-Alike Requirements: Certain datasets are made available with licensing terms that make them difficult to use in AI/ML. For example, “the notion of derivative work is ill defined” in the Creative Commons Share Alike license (CC-SA) regime.

(vi) Licensing Language Requires Standardization and Context-Appropriate Adaptations to ML and AI: [the use of this statement is unclear – see “Comment”, above]. 


The lack of clarity make it difficult for database users to determine if their intended use cases fall within the granted permissions, leading to unpredictability and the need for further analysis.

This in turn increases transaction costs, as resources, time, and expertise are required to assess whether a database can be used for a particular purpose. The paper aims to resolve these conceptual ambiguities using a constructed taxonomy that clarifies and standardizes data licensing terms.

===============================

Part (3) Taxonomy underlying the MDL (p.5-9)

The MDL approach is to create a Use Case Taxonomy, where each Use Case is a framework that uses defined legal rights to use data in ML and AI modeling . These are the definitions:

  • Data: Raw Data, including metadata, which provides structural details about the data.
  • Labelled Data: Data enriched with metadata labels or tags, which help models make sense of the data.
  • Model: Refers to machine learning (ML) or AI algorithms used to derive insights or predictions.
  • Untrained Model: The model which has not been exposed to data to optimize its parameters.
  • Trained Model: The model has been exposed to data which has been used to train the model.
  • Representation: A transformed version creates the means to use the model without touching the original data.
  • Output: The results generated by applying a trained model to data.

Application to a "Market Trading" use case
  1.  Evaluating Models: With this license, the licensee can train and test various versions of a model using the data to assess performance, but cannot use the output for stock trading or modify the model’s structure (e.g., keeping the trained model's weights).
  1. Research Use: This license allows the licensee to train models and create new datasets based on the data, with the restriction that any resulting models or datasets are subject to the same limitations. The license allows experimentation without commercializing the outcomes unless a separate paid license is negotiated.
  1. Publishing Research: This right permits the licensee to publish all research results, including those generated by models trained on the data, under the same restrictions as the original data. It clarifies the rights for both academic and private entities to advance the field of ML/AI and addresses ambiguity about commercialization rights in academic contexts.
  1. Internal Use: This license allows the licensee to train models and use them for trading proprietary capital for profit, but restricts sharing or selling the predictions to third parties.
  1. Output Commercialization: With this right, the licensee can commercialize the output by offering a service like a stock prediction API or trading third-party capital, but cannot sell or distribute the model itself.
  1. Model Commercialization: This license allows the licensee to both commercialize the model and its output, including selling the model itself, such as offering perpetual licenses for customers to modify and distribute the model.

Possible Restrictions
  • Designated Third Parties: Limiting data use to specific entities.
  • Sub-licensing: Restricting sublicensing the data to third parties, preventing contractors from using the data.
  • Attribution / Confidentiality: Licensors could require or prevent attribution.
  • Ethical Considerations: Ethical clauses could limit the data's use in certain fields, such as healthcare or military applications, to address concerns over the impact of data use.


Application 

The definitions are used to explain treatment of a database of historical equities trades.

===============================

Part (4) Caveats and Examples (p.9-10)

Licensors may include additional restrictions; rhe authors notinf that data licensing for AI/ML could be subject to external rule-making and controls, such as:

  • Legal Frameworks and Copyright: The underlying data may be subject to rights that are distinct from the licensing of use of the data.
  • Ill-acquired Data: There may be issues such as violations of privacy and owners' rights. 
  • No Property Rights in Data: The paper does not propose that data should be treated as property with inherent rights.
  • Database Rights vs. Copyright: "The paper distinguishes between specific database rights (e.g., the EU’s Database Directive) and copyright statutes (e.g., in Canada), noting that its framework focuses on contracts for accessing databases rather than recognizing specific legal statuses for databases." 

===============================

Parts (5 & 6) Conclusion & References (p.10-11)

This section outlines the resources and tools provided in the paper to promote clearer data licensing for AI and ML; and a list of references is included.

===============================

Appendices (p.12-16)

  • Appendix 1: Overview of commonly used datasets
  • Appendix 2: Summary of rights granted in conjunction with Models
  • Appendix 3: Top Sheet for Licensed Rights

  • Appendix 4: CC-BY4 Montreal Data License (MDL) attribution notice & cautions

09 November 2024

Policy Brief: What is the Value of Data? - Coyle & Manley (2022)

Overview: This brief explores various methods for valuing data, acknowledging limitations in existing models and emphasizing the need for more comprehensive approaches. The Typology of valuation methods in Part (2) below (and page 4 in brief) is especially helpful.

Comment: The global economy has been reshaped by data, with data-driven firms becoming dominant market leaders. This transformation extends to both private and public sectors. Although data’s value is widely acknowledged, there remains no consensus on how to quantify that value. This lack of clarity hinders optimal investments and governance.

This report builds on prior work (reviewed earlier in this blog - Ed) at the Bennett Institute for Public Policy (Coyle et al 2020). This article covers development of new methods of data value measurement.

Coyle, D. and A. Manley, Policy Brief: What is the Value of Data? A review of empirical methods, July 2022, Bennett Institute for Public Policy, University of Cambridge, July 2022: https://www.bennettinstitute.cam.ac.uk/wp-content/uploads/2022/07/policy-brief_what-is-the-value-of-data.pdf

  • Part (1) Discusses the value of economically valuable data to public and private sectors
  • Part (2) Reviews proposed data valuation methodologies
  • Part (3) Presents a framework to develop an estimate of value

===============================

Part (1) Introduction

In recent years, data has become a key driver of economic transformation, with data-driven companies making up seven of the top 10 firms globally by market capitalization in 2021. This shift is particularly evident in the growing productivity and profitability gap between data-intensive firms and others.

Data’s value is increasingly recognized across sectors, including the public sector. Despite this recognition, no consensus has developed to empirically measure the value of data, hindering its full potential. While many firms and investors acknowledge the value of data, particularly through data services and stock market evaluations, the absence of clear valuation methods makes it difficult to guide investment decisions or govern data usage effectively. Coyle & Manley's report explores various approaches to data valuation; and highlights the challenges of incorporating factors like opportunity costs, risks, and the costs associated with data collection and storage into such assessments.

===============================

Part (2) Proposed Data Valuation Methodologies

Coyle (2020) presented the Lens framework, describing data through the Economic Lens and the Information Lens. 

Building on prior research, the report identifies key characteristics influencing data's value and outlines valuation approaches. Traditional methods—cost-based, income-based, and market-based—are commonly used but fail to fully account for other inputs. 

The report also highlights newer approaches which improve capture of data's broader economic value, such as ascribing value using data flows and marketplaces. Comparative methods are cited, including:

• Internet of Water: A taxonomy of various data valuation methods (2018).

• Ker and Mazzini (2020) identified four different methods: (a) cost-based; (b) income-based approach; (c) market capitalisation; and (c) trade flows.

• OECD Going Digital Toolkit: Estimates the value of data, summarizes the System of National Accounts cost-based frameworks being adopted by governments; and summarized other approaches.

The methods are summarized as:

  • 2.1 Cost-based Methods
  • 2.2 Income-based Methods
  • 2.3 Market-based methods (Marketplaces, Market capitalisation, Data Flows) 
  • 2.4 "Ambiguity-driven" methods (term coined by blogger - Ed)
  • 2.5 Impact-Based Methods

-----------------------------

2.1 Cost-based Methods

These methods calculate the costs of generating, storing, and replacing data, providing a lower-bound estimate of its value. Variants like the Modified Historical Cost Method (MHCM) adjust for data characteristics, and the consumption-based method reflects usage rates. National statistical offices, such as Statistics Canada and the UK Office for National Statistics, have trialed this method. Cost-based methods are widely used for valuing data, rooted in the System of National Accounts (SNA) (1).

[Government interest is to define the value to generate taxes - Ed]. Their challenge is that "national level cost-based approaches rely on having well-classified data at the microlevel. This will be difficult to achieve and there are several blurred lines that make classification harder." (p.5-7) 

2.2 Income-based Methods

These methods estimate data's value through expected revenue streams generated by the data, such as selling marketing analytics. A common approach is the "relief from royalty" method, which estimates savings from owning data rather than licensing it. However, challenges arise in distinguishing data's contribution to revenue, especially for firms where data enhances products rather than being sold directly. This method also introduces uncertainty as it relies on judgment.

2.3 Market-based methods

These methods use observable prices for data, though such prices are rare since most data is used internally. When available, market prices offer valuable insights but reflect only a partial estimate of the broader social value of data. Key academic approaches include using data marketplaces, market capitalization of firms, and global data flows to estimate value. However, limitations remain, especially when data is aggregated or traded in complex ecosystems like credit scoring, where the true value often exceeds the sum of its parts.

2.3.1: Data Marketplaces

The literature on data marketplaces explores their potential to increase the value of data by reducing transaction costs, improving pricing transparency, and allowing multiple users to derive value from the same datasets. However, the success of such initiatives has been inconsistent, with key challenges including complex pricing mechanisms, regulatory differences, and a lack of trust. Data suppliers often bundle datasets and set prices based on consumer willingness to pay, but much of the literature remains theoretical and idealistic. Case studies from China, New Zealand, the EU, and Colombia show that effective data pricing and trust in data quality are critical for success, yet low participation often undermines marketplace efforts. Historical examples, such as Microsoft's failed Azure DataMarket, highlight the difficulty in building customer interest, while current platforms like the Shanghai Data Exchange and Ocean Market demonstrate varying approaches to data transaction and pricing. Ultimately, while data marketplaces hold significant potential, barriers such as trust, pricing, and regulation continue to limit their effectiveness.

2.3.2 Market capitalization-based

These methods are used to value data by examining its impact on a firm's market value, particularly for data-driven companies. This approach estimates the worth of these firms by looking at their overall market capitalization, which includes the value derived from data and analytics. For example, Ker and Mazzini (2020) estimate that U.S. data-driven firms, identified through lists like "The Cloud 100," are collectively worth over $5 trillion. Coyle and Li (2021) further build on this by analyzing how data-driven companies, such as Airbnb, disrupt traditional firms like Marriott, leading to the depreciation of incumbents' organizational capital. The decline in the value of non-digital firms' organizational capital due to data-driven competition helps estimate how much these firms should be willing to pay for data. Overall, these methods provide a way to quantify the value of data in relation to firm competitiveness and valuation in data-intensive industries.

2.3.3: Data Flows

Data flows are valuable in markets with limited information, as they can be observed and analyzed, with a strong correlation between the volume of data flow and its value on dominant online platforms. However, quantifying global data flows is challenging because most assessments only account for data that crosses international borders. While Ker and Mazzini (2020) and Coyle and Li (2021) suggest that the link between data flow volume and data value in specific locations is weak, due to large data hubs serving broad areas and the need for local knowledge, they highlight the economic significance of the content in data flows over volume alone. For example, video streaming generates more traffic than e-commerce but contributes less economic value. Their research also underscores the economic importance of digitally deliverable products, noting that many countries still lack a framework for categorizing digital trade. Overall, while data flows help understand data's economic value, their measurement is still evolving, hindered by definitional and geographical complexities. 

-----------------------------

Section 2.4: Ambiguity Methods: 

This section discusses Experiments and Surveys to estimate value where market prices do not exist. 

-----------------------------

2.5: Impact-Based Methods

The intent of data collection is to develop stories from which to mine insights for decision-making. Impact-based methods assess the value of data using cause-and-effect measurement. These create greater value-add, "making their value more persuasive than traditional quantitative approaches." These methods use testing, such as comparative scenarios, to test response. Slotin (2018) reviews five data valuation methodologies, favoring impact-based approaches for their clarity and communicative strength. These include:

(a) Empirical Studies: Arrieta-Ibarra et al. (2020) estimated that "data use accounts for up to 47% of Uber’s revenue. In a scenario where drivers are fully compensated for their data, this could equate to $30 per driver per day for data generation."

(b) Decision-Based Valuation: A variant of empirical-based studies, this method "adjusts the value of data based on factors like frequency, accuracy, and quality before weighing outcomes by their contribution to decisions. This method acknowledges that value derives from improved decision-making and considers alternatives to using the data, although it requires subjective judgment."

(c) Shapley values: This variant ignores the determination of value-add of insights. Instead, Shapley values represent a subset of impact-based methods that focus on valuing data in its raw form rather than solely based on its applications in data-driven insights. Originating from game theory, Shapley values provide a unique payoff solution within a public good game, ensuring group rationality, fairness, and additivity. 

Shapley values are used in computer science to evaluate the contribution of individual data points to model performance, assess data quality, and optimize feature selection. They are also applied to determine compensation for data providers by quantifying the value of their data.  Although Shapley values provide a useful framework for valuing data, they represent only one possible solution to public good problems, with alternative approaches that might offer better properties. The method has advantages, such as identifying valuable data for collection, but also faces drawbacks, including high computational costs and challenges in translating value into monetary terms. 

(d) Direct measurable economic impact: Various studies analyze the growth and jobs impact.

(e) Stakeholder-based methods: These methods analyze the value to the sector supply chain of data availability: "This is a wider definition of value and may include value upstream or downstream omitted from other methods; it can encompass the non-rival aspect of data. The data consultancy Anmut is one that has developed this method and provides a case study of their valuation of Highways England data (Anmut n.d.). The challenge of this approach is that it requires professional judgment of value vs. auditable measurements of value.

(f) Real options analysis: Real options analysis provides a method for estimating the value of data by considering its potential future use cases rather than its current applications. This flexibility enables firms to capitalize on positive opportunities while minimizing downside risks. Data is considered non-rival, meaning its value does not diminish with use, and its potential applications can remain undefined at the time of collection. The option value represents the "right but not the obligation" to generate insights from data in the future, allowing firms to collect data for unknown future purposes. The value here is that this incentivizes waiting to determine data value in advance; by assessing the impact of more new information (policy changes, tech changes, shifts in consumer preferences) before deciding whether to analyze the data. 

===============================

3. Discussion

Policymakers are stuck because there is no "consensus or best method for valuing data." The authors instead propose setting up a schema to classify methods and to validate them with external surveys. The schema goal is to determine: 

  1. What is being valued?
  2. Who is valuing the data?
  3. When is the valuation taking place?
  4. What is the purpose of the valuation?
-----------------------------

3.1 What is being valued?

There are a variety of things that could be referred to as ‘data.’ The possible distinctions are illustrated in the ‘data value chain,’ (see blog review of Coyle, 2020) which sets out different stages from the generation of raw data up to the decisions made using data insights generating the potential end-user value. The authors note that: "In general, the raw data is of least interest, and some of the literature goes as far as to state that raw data does not hold any value on its own...(as) even with cost-based methods, in many ways the most straightforward approach, it is almost impossible to distinguish between costs associated with raw data generation and database formation." 

[Note: OrbMB's ORBintel method illuminates the costs and we forecast developing the means to establish the value of raw data that is to be collected, in advance of the need - Ed.]

-----------------------------

3.2 Who is Valuing the Data? 

This section explores how the perspective of different stakeholders affects the methods used to value data (Data Producers, Private Sector Producers, Data Users, Data Hubs, Public Sector vs. Private Sector Valuation, Intangible Asset and Productivity).

The key point is that Data serves as an intangible asset that provides firms with a productivity advantage, especially when coupled with complementary skills. Firms that capture monopoly rents from data tend to value it higher than alternative users, leading to a divergence between private and social valuations of data.

Valuation varies significantly based on the perspectives of different stakeholders, with public sector approaches emphasizing societal value and the monopoly rent that is taxation, while private sector valuations focus on costs and impacts. Understanding these differing perspectives is crucial for accurately assessing data's overall worth.

-----------------------------

3.3: When is the Valuation Taking Place? 

This section distinguishes between ex ante (before the event) and ex post (after the event) data valuation methods: Ex Ante Valuations are fraught with uncertainty, therefore risk and cost; therefore less widely used. Ex Post Valuations mitigate the uncertainties.

-----------------------------

3.4 What is the purpose of the valuation? 

The authors contend that the purpose influences the choice of methodologies, with different approaches suited to different goals. 

===============================

References:

(1) SNA is the United Nations (UN) framework setting “the internationally agreed standard set of recommendations on how to compile measures of economic activity.” The framework establishes consistent accounting rules and classifications, to make multistate comparison possible. The next update (2025) will include an update on data valuation. 



09 October 2024

More than Dual-Use?

We've accepted an invitation to join the Peachscore+Gust Data-driven Accelerator. This month's post was to be a review of Cambridge researcher Diane Coyle's (2022) valuation analysis; and we are juggling. So, for this month, a short post about an observation about market sectors.

The Root Taxonomy of Goods and Services - More than Dual-Use?

Goods and services generally get classed into three sector classes: 

  • Civilian
  • Military/national security
  • Dual-Use (use cases in both sectors)

This has never been really exactly true, as there are numerous goods and services that are not purely civilian consumer (individual & household) offers. These include various public administration and safety segments ranging from police and rescue to wastewater and pothole maintenance.

The better structure might be to say that there are five sector classes:

  1. Civilian;
  2. Civil Aid; and 
  3. Military/National Security:
where:
  • #1, #2, and #3 are (#4) Triple-Use; and
  • #2 and #3 are (#5) Dual-Use:

Consider the carabiner: 


Comments?





28 August 2024

The Value of Data – Policy Implications – Main report – Coyle (2020)

Overview: This is a fascinating contribution to the literature. The discussion about content ("Information Lens") driving utility is especially useful to understand where and how to determine when to apply cost/benefit analyses. The authors' proposed framework clearly expresses the need to incentivize private innovation to concurrently aid the public good - a partnership of private, non-profit and public stakeholders. 


Comment: The previous articles are concerned with the internal business value of data, and the economic value of data for statistical purposes. This article dives into the economic value of data to the State. The perspective is that of public sector economists discussing opportunities and concerns for the United Kingdom: noting that: "A greater understanding of the value of data would help identify where the benefits of greater investment in and sharing of data are worth the costs;" and that in this view, there is need for regulation “for establishing a trustworthy institutional framework for managing, monitoring and enforcing the terms of access.” 

Here, data is defined as intangible; and is generally viewed as a homogeneous economic good (Note 1); this leads to the view that valuation is not that important because market pricing sets the value.  

The authors note the absence of costing because “available empirical studies use market valuations or transactions” to estimate value. These authors note that “the value of different types…can be very different”; and value is tied to the use case (what is the utility?), which they contend leads to a need to determine utility (public, social, private) in which there is a public interest. Finally, the authors call for greater regulation to obtain greater “social welfare value” and to prevent private asymmetric information advantage. But recognize that regulation will affect ROI due to the need to collect and clean data, and to invest to develop complementary skills and assets. 

Note: The degree of not-sharing is often key to the profitable continuance of a private enterprise - Ed.

Coyle, D., S. Diepeveen, J. Wdowin, J. Tennison, and L. Kay, (2020) The Value of Data – Policy Implications – Main Report. Bennett Institute for Public Policy, University of Cambridge, and Open Data Institute, February 2020: https://www.bennettinstitute.cam.ac.uk/ publications/value-data-policy-implications/ (Accessed Q12024)

Next month, we review Coyle et al (2022). 

 ____________________ 

  • Part (1) Introduces the public policy interest in data value 
  • Part (2) Describes current valuation taxonomies and develops a two-lens framework, commencing with the Economic Lens 
  • Part (3) Deep dives into the second lens: the Information Lens 
  • Part (4) Provides an overview of the current UK legal framework 
  • Part (5) Discusses the features driving Market-based Valuation 
  • Parts (6,7,8) Delve into three related issues 
  • Part (9) Presents conclusions and recommends future stakeholder work plans 

____________________

Part (1) Introduction: The subject of this paper is to discuss the policy interest in data value; and to develop a schema to determine the value of data that is made available for public purposes that encompass all sectors of society.

 

(A)    The policy interest has two dimensions:

 

(a) Governments need to use data to make policy decisions; and

(b) Governments need to understand the value of the transaction of sharing (or not sharing) data, in the context of: 

 

(i) Understanding the value (and impact) of data transactions to government and society, and

(ii) the worrisome economic problem of private actors mining public data for private purposes at no cost (i.e. taking it for free to modify & resell or to restrict the sale of the improved resource to a limited clientele) [Similar to other State-resources that are licensed to extractors, who require return-on-investment to bear the cost of bringing the refined resource to market - ed.].

 

(B) The schema to determine data value consists of two lenses: Economic and Information (Note 2).

 

 

Part (2) Taxonomies and a new framework: This section describes current taxonomies of valuation; and then discusses part one of the new framework: the Economic Lens (The distinctive economic characteristics of data)

 

This section introduces the reader to the economic problem of valuing data, a resource which arises from "the creation of value from data of different kinds, its capture by different entities, and its distribution". There is a review of the intangible nature of the resource, high-level use cases, and the impact of externalities (impacts to data which "are often positive, such as additional data improving predictive accuracy, or enhancing the information content of other data": p.5).

 

Data is described as a non-rival asset (the same asset can be used by many users) which can be “closed, shared, or open. If access to data is restricted, its uses are limited; i.e. it becomes a private good. If it is shared with a select group of people - it is a club good - its uses and analysis can be wider, perhaps creating more value. If data is shared openly - a public good - anyone can use it.” [The economic actors who participate in the life of the State each can individually control data as a private good, group good, and public good. This includes the State, where-for example-records that are restricted are variously a private good, group good, and public good. - ed.]

 

The authors note that the unique ("intangible") nature of data gives it low utility if not shared, and therefore must be shared or licensed to make good use of the resource. The economic perspective expressed here is that data "is not best thought of as owned or exchanged;...that personal ‘ownership’ is an inappropriate concept for data (and that characterising data as ‘the new oil’ is similarly misleading)." [Note: this is for national statistical accounts and public services; “ownership” and “exchange” is how value is expressed by private buyers and sellers - ed].

 

The second part is to develop five new characteristics, to develop an improved valuation schema: (a) Marginal return; (b) Externalities; (c) Optionality; (d) Consequences; (e) Costs.

 

Part (3) discusses part two: the Information lens (the determination of 'economic utility' value):

 

The Information Lens is defined as the determination of utility to express economic value. Use cases are organized to fit five subject areas, and then five sub-frameworks:

  • (1a) People
  • (1b) Organization Type
  • (1c) Natural environment
  • (1d) Built environment; and
  • (1e) the type of goods or services being offered.

(2) Use Cases are organized into five sub-frameworks, and there is an example table:

  • (2a) Generality of the use case (as a non-rival good, data is repeatedly useful for many purposes; "Generality" is a determination of use case repeatability);
  • (2b) Granularity (this is the degree to which data is "filtered, aggregated and combined in different ways to reveal different insights" p.8)
  • (2c) Geo-spatial coverage: the area that data refers to limits or develops its utility.
  • (2d) Temporal coverage: data utility is time-bound;
  • (2e) Human User Groups: This sub-schema organizes human users into three groups: Planners (ex. city planner, Operators (commuters), and Historians (police investigating a crime).

The section continues with a review of technical characteristics that contribute to valuations. These include: Quality (p.10), Sensitivity and personal data (p.10), Interoperability and linkability (developing standards and common identifiers to improve aggregation) (p.10), Excludability (data that is not common and so much be actively shared (p.11), and Accessibility (the degree to which data dissemination is organized and executed) (p.11). 

 

Here, the Open Data Institute's Data Spectrum is charted to identify "some of the access conditions determining whether data is a private, shared, or public good. Access conditions can be determined by technology, licensing or terms and conditions, and regulation." (p.12).

Part (4) provides an overview of the current UK legal framework, which includes Intellectual property rights and licensing (p.14), Intellectual property rights in public sector information (p.15), and Data protection rights (p.16).

Part (5) dives into the features driving Market-based valuation of data.

 

The primary methods are to use Market Price methods to estimate value:

 

(a) Stock market valuations;

(b) Income-based;

(c) Cost-based valuation methods.

 

Method (a) is to compare data-driven companies and non-data-driven companies; analyses suggest that the former have become more valuable that the latter.

 

Method (b) is to make an estimate of future cash flows (Free Cash Flow or FCF) derived from the asset (Note 3) where the "data value chain" has been developed to visualize this method. The FCF method can be successfully exploited by e-commerce companies such as Amazon that enjoy feedback loops; where insights drawn from data help improve the customer experience, which improves per customer unit sales. Note: Many companies in the space obtain customer data as a free good (not paid for).

 

Method (c): The public sector's cost approach is to estimate "the aggregate value of data to the economy in the national accounts, as there are relatively few market sales of datasets, with most being generated within the business in the process of providing other goods and services."

Parts (6,7,8) Delve into existing non-market estimates, Creating value through open and shared data, and Institutions for the data economy.

 

The authors note that: "Market valuations thus provide useful information but do not capture the full social value of data." Their contention is that-- subject to trade-offs' analysis--it would be economically better to create value through open data sharing. Here, citing O'Neill, the authors note that "the real or perceived crisis of trust in many societies reflects suspicion of authority" which—to this reader—suggests that the survival of a system of government that is trusted requires honesty and integrity in data management per "the social and legal ‘permissions’" of the society.   

 

Trade-offs to consider are the need to:

  • incentivize investment and innovation,
  • maintain data security, and the related need to
  • protect personally and commercially sensitive data. 

It is hard to develop and adhere to this form of social navigation. The authors suggest three frameworks:

  • Elinor Ostrom’s framework for the management of shared resources: [chart: Ostrom’s principles]
  • Paying new attention to creating proper Data Infrastructure (vs. leaving it all to private actors);
  • Establishing data trusts, data pools, and brokerages (to create a level playing field of data access).

Part (9) presents conclusions and future work plans.  

The authors note that “the quality of data is a key challenge” in-and-of-itself; and that the “quality needed depends on what the data is being used for.” The authors also state that: “Asymmetries of information mean that contracts for data use are incomplete, and the regulatory framework should recognise this, particularly that schemes for sharing data in a regulated way change the returns on investment in collecting and cleaning data, and investing in complementary skills and assets.” In this context, the authors express the need to develop different use cases to determine utility (public, social, private) in which there is a public interest; whilst specifying the need to incentivize private investment to efficiently deliver data goods and services for every sector of society. The proposals here are to:

 

·         Incentivise investment without disincentivising sharing

·         Limit exclusive access to public sector data

·         Use competition policy to distribute value

·         Explore mandating access to private sector data

·         Provide a trustworthy institutional and regulatory environment

·         Simplify data regulation and licensing

·         Monitor impacts and iterate

----------------

(1) The definition is derived from the UN System of National Accounts or SNA.

 (2) "Lens analysis requires you to distill a concept, theory, method or claim from a text (i.e. the “lens”) and then use it to interpret, analyze, or explore something else" cf.: https://pressbooks.cuny.edu/qcenglish130writingguides/chapter/lens-analysis/

 
(3) cf. 26 C. Mawer, “Valuing data is hard,” Silicon Valley Data science blog post (2015). Accessed by the authors at: https://www.svds.com/valuing-data-is-hard/. See also C. Corrado, “Data as an Asset,” presentation at EMAEE 2019 Conference on the Economics, Governance and Management of AI, Robots and Digital Transformation, 2019; and M. Savona, “The Value of Data: Towards a Framework to Redistribute It,” SPRU Working Paper 2019-21 (October 2019).
 
(4) cf: O. O’Neill, “A Question of Trust,” BBC Radio 4 (2002). Accessed by the authors at: www.bbc.co.uk/radio4/reith2002/


The Land of Startups is a Land of Pivots

We've graduated from the Peachscore+Dealum Accelerator , and in the midst of another pivot; having identified a beachhead industry for o...