New in mid September: OpenAI and Harvard released the largest study yet of ChatGPT consumer use. An NBER paper on how to count AI argues the value of data rises with the value of AI services. Both point to the same conclusion. Models need a continuous flow of live human signal data. That is why Reddit-class communities and RSL-style licensing move to the center of the AI economy.
Download OpenAI x Harvard paper
Download NBER w34330
Thesis. AI does not just learn language. It learns what matters to people right now. The usage study shows the mass market leans on chatbots for guidance, judgment, and writing. The economics paper frames data as an input whose value rises with the outputs it powers. The scarce resource is shifting from raw text to consented human context, and the revenue model is shifting from scraping to licensed access.
Primary sources:
- NBER Working Paper by OpenAI and Harvard: How People Use ChatGPT, Sept 2025
- NBER Working Paper: Making AI Count: The Next Measurement Frontier, Oct 2025
Recent evidence, at a glance
Signal
What changed
Why it matters for models
Market implication
OpenAI usage
About 700M weekly active users by late 2025. 2.63B daily messages in June 2025. Topic mix: Practical Guidance 28.8%, Seeking Information 24.4%, Writing 23.9%.
Users want advice, tone, and judgment. That needs live human norms, not static facts.
Perpetual demand for behavior-labeled and time-stamped discourse.
Economic valuation
NBER argues data’s value is circular. The value of inputs rises with the value of AI outputs.
As AI gets better, the best human data becomes more valuable.
Supports recurring licensing and data-as-capital pricing.
Dynamic vs static data
Recent research shows refreshed training data can outperform frozen sets at the same size, with meaningful accuracy gains.
Self-training on your own outputs invites feedback loops and style drift.
External human signal remains a strategic dependency.

OpenAI and Harvard, Figure 3. ChatGPT consumer WAU grew to about 700M by September 2025.

OpenAI and Harvard, Table 1. Daily messages reached 2.63B in June 2025. About 73 percent were non-work.

OpenAI and Harvard, Figure 7. Guidance plus information plus writing accounts for roughly 77 to 78 percent of use.
What dynamic human signal data means
It is the live, messy layer of the internet. What people ask, approve, reject, and debate, with timing and engagement baked in. Subreddits are labeled communities of behavior. Moral judgment in r/AITA. Risk language in r/WallStreetBets. Empathy in r/relationship_advice. Expert challenge in r/AskScience. That structure creates training signals you cannot fake with static text.
Where Reddit fits and why licensing wins
Subreddit example
Signal type
Model capability
Why value persists
r/AskReddit
Collective judgment and cultural norms
Dialogue tone and value alignment
Endless edge cases and vote-based labels
r/AITA and r/relationships
Moral dilemmas and empathy
Safety tuning and harm analysis
Nuance that no textbook captures
r/WallStreetBets
Sentiment under volatility
Behavioral finance reasoning
Time-stamped outcomes for labeled supervision
r/AskScience and r/technology
Expert Q and A and falsification
Factuality calibration
Peer correction plus taxonomy
The economic turn from scraping to consent
RSL lets publishers declare machine readable terms for AI agents. Who may crawl, at what price, and under what license. Combined with CDN checks and license servers, you get a practical gate. Agents present tokens. Access is metered. Usage is auditable. The NBER w34330 paper adds the frame. Data’s value is circular. As AI services get better, the input data that powers them becomes more valuable. That is a structural case for recurring payments to the sources of live human signal.
Data source
Temporal relevance
Context depth
Replaceable
Control
Books and Wikipedia
Low
Medium
Yes
Open corpus
News archives
Medium
Medium
Partial
Publishers
Reddit and forums
High
High
No
Platforms via RSL and API
Product telemetry
High
Medium
Feedback loop risk
Labs
Synthetic and self play
Low to Medium
Narrow
Limited
Labs
Findings
- Usage is the tell. Guidance and writing dominate consumer demand. Those rely on current human norms and tone.
- Data’s value is reflexive. As AI output value rises, the best input data appreciates. That supports licensing, not scraping.
- Dynamic beats static. Refreshing training data outperforms frozen sets at the same size.
- Reddit is a durable supplier. Labeled and time stamped human discourse is scarce and compounding in value.
Investor Takeaway
- From compute to consent. The scarce input is live, licensed human signal. Platforms that aggregate it gain pricing power.
- Recurring rails. RSL and license servers turn discourse into metered and auditable cash flows.
- Risk to closed loops. Heavy self training invites feedback drift. Diverse external data is the hedge.
- Valuation shift. Treat high quality human signal like capital with compounding yield.
Bottom Line
AI stays smart by staying human. The latest evidence says models need consented, behavior labeled discourse on a continuous basis. That makes Reddit class communities and RSL class rails the quiet utilities of the next decade.
Disclaimer: This content is for informational and educational purposes only and does not constitute investment advice or a recommendation to buy or sell any security. Third party data may contain inaccuracies. Always conduct independent diligence and consult licensed professionals.