methodology

/memory-loss/methodology · config 2026-09-19.1

Estimating global data creation

Memory Loss sums annual estimates of the global datasphere from January 1, 1999, then projects them forward second by second. It is a model of data created, captured, copied and consumed—not a live measurement, a count of unique internet content, or a count of what still exists.

01 The headline number

All visitors with the same clock see the same total. The counter is evaluated from a UTC timestamp and the configuration; elapsed browser frames never accumulate into it.

Each year contributes its configured volume. Within a year, the rate grows exponentially, normalised so that the integral equals that year’s total.

CY(τ) = AY × (kτ/T − 1) / (k − 1)

rY(τ) = AY × ln(k) × kτ/T / (T × (k − 1))

A is the annual total, k is the following year’s total divided by this year’s, T is the number of seconds in the year, and τ is the time elapsed. When k = 1, the formula becomes linear. The headline adds the completed years to the current year’s integral. Leap days are included; leap seconds are ignored.

Missing years are interpolated geometrically between known totals. After 2029, the final known annual ratio is extended forward. The cumulative total is continuous at year boundaries; the rate can change there because adjacent forecasts imply different growth curves.

Annual data created · decimal zettabytes
yearZBbasis
19990.008approximate
20000.012approximate
20010.02approximate
20020.03approximate
20030.05approximate
20040.08approximate
20050.13approximate
20060.16approximate
20070.28approximate
20080.49approximate
20090.8approximate
20102approximate
20115approximate
20126.5approximate
20139approximate
201412.5approximate
201515.5estimate
201618estimate
201726estimate
201833estimate
201941estimate
202064.2estimate
202179estimate
202297estimate
2023120estimate
2024149estimate
2025181forecast
2026221forecast
2027295.083interpolated
2028394forecast
2029527.5forecast

The 1999–2014 values are approximate and contribute 3.66% of the launch total. The earliest years are back-extrapolations, not independent measurements.

Statista series, reproduced by DemandSage ↗
Background for the 2015–2026 series. The current 2024 figure is 147 ZB; this model version uses 149 ZB.

IDC: The Digitization of the World ↗
Primary background for the global datasphere definition. This older forecast is not a primary source for the complete 2026 config.

DesignRush: daily data generation ↗
Secondary compilation for later forecasts. Forecast editions differ.

Reading the launch values

At 2026-09-19T00:00:00Z, this configuration gives 1.012156 YB (1.012156e+24 bytes) and 7.43 PB per second. One yottabyte is 1024 bytes. The unit label uses this decimal definition throughout.

The full integer is evaluated with 36-place fixed-point arithmetic and BigInt. Those digits describe the projection; their length does not imply byte-level observational accuracy.

02 The categories

Category counters begin at 2026-09-19 with a zero offset unless an offset is configured. Their byte totals use assumed average file sizes. Six newer categories use dated published activity figures held flat; their AI attribution is unknown, so they are excluded from the origin calculation. They overlap, are illustrative, and do not add up to the headline.

Baseline rates and model assumptions · 2026-09-19.1
category / sourceitems / sannual growthAI weight
hours of video uploaded to youtubeSocialRails upload statistics ↗8.333+0%0%
photos & videos posted to instagramEarthWeb data statistics ↗1100+5%0%
instagram storiesBusinesstats: content per minute ↗11583+5%0%
snaps sentSnap figures via SocialRails ↗63657+5%0%
whatsapp messagesBusinesstats: content per minute ↗683333+5%0%
emails sentThe Radicati Group ↗4351852+3%0%
messages sent to ai assistantsNBER working paper 34255 ↗46300+100%100%
ai images generatedWearView: AI image statistics ↗1389+100%100%
ai videos generatedAdwave: AI video generation statistics ↗58+150%100%
songs uploaded to streamingMusic Ally / Luminate ↗1.74+20%50%
ai songs generatedBillboard: streaming upload report ↗81+100%100%
posts uploaded to tiktokSteel et al.: Just Another Hour on TikTok (v5, 2026) ↗3116.898+0%unknown
podcast episodes publishedListen Notes: podcast statistics dataset ↗0.863+0%unknown
wordpress blog postsWordPress.com: publishing activity ↗26.618+0%unknown
reddit posts + commentsReddit: July–December 2025 transparency report ↗139.523+0%unknown
code commits pushed to githubGitHub: 986 million commits in its 2025 report ↗31.266+0%unknown
hours broadcast on twitchStreamlabs / Stream Hatchet: Q4 2025 report ↗26.092+0%unknown
hours of video uploaded to youtube · source note

500 uploaded hours per minute; rounded to 8.333 per second. Secondary compilation; the flat rate is a model assumption.

photos & videos posted to instagram · source note

95 million posts per day; rounded to 1,100 per second. Historical secondary estimate. Includes video posts, so the picture comparison is illustrative.

instagram stories · source note

695,000 Stories per minute. Secondary estimate; not a live platform count.

snaps sent · source note

5.5 billion Snaps per day. Snap newsroom figure cited by a secondary compilation.

whatsapp messages · source note

41 million messages per minute. Secondary estimate; messages use a 2 KB average.

emails sent · source note

Approximately 376 billion emails per day. Approximate vendor projection. The exact report edition has not been independently verified.

messages sent to ai assistants · source note

ChatGPT 2.5 billion messages per day, multiplied by 1.6 for other assistants; approximately 4 billion per day. The cross-platform multiplier and 100% annual growth are model assumptions. The underlying paper has not been independently verified for this model version.

ai images generated · source note

120 million images per day; an assumed midpoint within an 80–200 million range. Model estimate informed by platform disclosures and secondary compilations; not a measured cross-platform total.

ai videos generated · source note

Approximately 5 million generated video clips per day. Model estimate informed by Veo/Flow and Kling figures; clips vary substantially in duration and size.

songs uploaded to streaming · source note

150,000 tracks uploaded per day, rounded to 1.74 per second. September 2026 report has not been independently verified for this model version. A 50% AI weighting is an assumption, not a verified industry-wide share.

ai songs generated · source note

Approximately 7 million generated songs per day, based on a reported Suno figure. Generated songs and streaming uploads may overlap. This is an estimate, not a count of unique released songs.

posts uploaded to tiktok · source note

269.3 million posts over the sampled day, April 10, 2024; divided by 86,400 seconds. Historical one-day research estimate, carried forward flat. Version 5 revises earlier estimates. Includes posts later unavailable. The current platform rate and AI share are unknown.

podcast episodes published · source note

27,214,183 new episodes in 2025 in the dataset checked September 19, 2026; divided by 31,536,000 seconds. Public RSS podcasts indexed by Listen Notes, not every audio or video podcast. The dataset is revised over time and filters some AI/spam content. Held flat; AI share is unknown.

wordpress blog posts · source note

70 million new posts per month; multiplied by 12 and divided by an average 365.25-day year. WordPress.com and connected Jetpack network only, not all blogs or web pages. The page does not date this long-running estimate. Held flat; AI share is unknown.

reddit posts + comments · source note

4.4 billion posts and comments during 2025; divided by 31,536,000 seconds. Posts and comments combined, excluding private chats. Held flat. Text-byte estimate excludes attached media; AI share is unknown.

code commits pushed to github · source note

986 million commits reported for the 2025 reporting year, normalised to 365 days. A commit is a change, not a unique file or software release. GitHub only; may include forks, automation and AI-assisted work. Held flat; AI share is unknown.

hours broadcast on twitch · source note

207.4 million creator-hours broadcast on Twitch in Q4 2025; divided by 92 days. Counts hours broadcast, not audience-hours watched. Held flat. Byte estimate represents one encoded stream, excluding viewer copies and multiple renditions; AI share is unknown.

Average sizes

Byte assumptions used for category totals
itembytes / itembasis
hours of video uploaded to youtube2 250 000 000One hour of 1080p video at 5 Mbps.
photos & videos posted to instagram3 000 000Typical 12 MP HEIC/JPEG; applied to a mixed post category.
instagram stories4 000 000Mixed photos and short videos.
snaps sent4 000 000Mixed photos and short videos.
whatsapp messages2 000A text message; attachments averaged into a simple 2 KB assumption.
emails sent75 000Body and headers, with attachments amortised.
messages sent to ai assistants4 000A text-only prompt and reply pair.
ai images generated1 500 000A 1024 × 1024 PNG/WebP image.
ai videos generated15 000 000An 8-second clip at 1080p.
songs uploaded to streaming9 000 000A 4-minute track at 320 kbps, rounded to 9 MB.
ai songs generated9 000 000A 4-minute track at 320 kbps, rounded to 9 MB.
posts uploaded to tiktok10 210 000Assumed 4 Mbps × 20.42 seconds ÷ 8 = 10.21 MB. Duration comes from the sampled posts; bitrate is an illustrative assumption.
podcast episodes published43 200 000Assumed 45-minute audio episode at 128 kbps = 43.2 MB. Duration and bitrate are illustrative, not measured averages.
wordpress blog posts100 000Assumed 100 KB of text and markup per post, excluding embedded media and page assets.
reddit posts + comments2 000Assumed 2 KB per post/comment for text and metadata, excluding attached images and video.
code commits pushed to github50 000Assumed 50 KB per code change, not an entire repository or a measured average diff size.
hours broadcast on twitch2 250 000 000Assumed 5 Mbps × 3,600 seconds ÷ 8 = 2.25 GB per creator-hour; one encoded stream.

03 AI-generated and other data

AI share = Σ(rate × bytes/item × AI weight)
÷ Σ(rate × bytes/item)

AI images, videos, songs and assistant exchanges have a machine weight of 1. Streaming uploads have a weight of 0.5; the original other-data categories have a weight of 0. New categories with unknown attribution have no weight and are excluded from both the numerator and denominator. “Other” is the remaining model bucket, not verified human authorship: ordinary messages and posts may also be automated.

A reported Deezer example informed the 50/50 streaming assumption, but one service’s share is not proof of an industry-wide split. AI assistant exchanges include a human prompt. Generated songs can later be uploaded, so the categories are not mutually exclusive.

At launch, this model yields 0.5920% AI-generated by bytes across the selected categories. It does not estimate the machine share of the entire global datasphere.

The picture readout divides the combined rate of Instagram posts, Stories and Snaps by the AI-image rate: 54.96 platform items per AI image at launch. Because some platform items are videos and can overlap, this is not a measured fraction of unique new pictures.

The song readout divides AI-song generation by all streaming uploads: 46.55 at launch. It says “streaming upload” because the denominator includes the configured AI share; describing it as a human-only upload would be incorrect.

04 Growth assumptions

Annual growth rates are modelling choices, not sourced predictions. The baseline is 2026-09-19. Most human categories grow slowly; YouTube upload hours remain flat. The initial doubling of several AI rates represents a period of rapid adoption, with greater uncertainty than the headline series.

For any category with a growth multiplier above 1.5, the instantaneous annual multiplier decays linearly toward 1.2 over 5 years. It then stays at that floor. This keeps totals increasing while avoiding indefinite doubling.

g(t) = 1.2 + (g₀ − 1.2) × max(0, 1 − t / 5Y)

rate(t) = rate₀ × exp(∫₀ᵗ ln(g(u)) du / Y)

items(t) = offset + ∫₀ᵗ rate(u) du

Y is 31,557,600 seconds. The logarithmic growth integral is evaluated analytically. Item totals use 64-slice Simpson integration within the decay period, then an exact exponential tail. Categories without decay use a closed-form integral. Times before the baseline clamp to the baseline.

Values and origin ratios in a frame use the same timestamp. A bounded cache reduces repeated work; it does not accumulate totals. Hidden tabs stop rendering and recompute from the wall clock when visible again.

05 Units

Explore the visual Units of Measure guide ↗

Decimal is the default: 1 TB = 1012 bytes, 1 YB = 1024 bytes. Binary uses powers of 1,024: 1 TiB = 240 bytes. This preference changes compact labels; the full byte integer and the model stay the same.

display units

Saved on this device when browser storage is available.

06 Known gaps

The following forms of media are visible on the counter page without invented totals. “No reliable estimate” means this model has no verified, comparable creation rate; it does not mean zero activity.

Facebook, X, Threads + other social posts

No current, comparable publishing rate has been verified for this model. Platform audiences and impressions do not measure posts created.

documents, PDFs, spreadsheets + slides

Creation is spread across private devices, workplaces and cloud services. No defensible global creation rate is included.

voice messages + video calls

These create audio and video, but participant-minutes, recordings and transmitted copies are different quantities. No comparable creation-rate estimate is included.

photos + videos kept on devices

Social uploads cover only part of photography and video creation. A global rate for files that remain on personal devices has not been verified.

ebooks + audiobooks

ISBNs, editions, formats and retailer listings overlap. No global count of newly created digital files is included.

games, animation + 3D assets

Releases, updates, textures, models and user-created worlds have very different sizes and overlap. No reliable combined creation rate is included.

  • The global datasphere includes offline, enterprise and device data, as well as copies and consumption. It is broader than the internet.
  • TikTok uses a historical one-day research sample from April 2024, not a current platform disclosure. WordPress covers only its publishing network. Neither represents all social media or web pages.
  • Email is a vendor projection. File sizes vary enormously. Category inputs mix dates, reporting scopes and secondary estimates.
  • This version uses 149 ZB for 2024; the linked secondary table currently says 147 ZB. That discrepancy is retained and disclosed so the version remains reproducible.
  • Some linked platform and 2026 media claims have not been independently verified for this model version. Source notes identify these limitations; a working link does not validate a claim.
  • Forecasts are uncertain, especially far beyond the last refresh. The counter depends on the device’s clock. A clock adjustment can move the displayed total backward.

07 Change log

2026-09-19.1
Added TikTok posts, podcast episodes, WordPress posts, Reddit discussions, GitHub commits and Twitch broadcasts, with dated sources and explicit size assumptions. Unmeasured media are listed as gaps. Unknown origins are excluded from the AI-share denominator. Headline and original category baselines are unchanged.

2026-09-19
Initial model: annual datasphere estimates, eleven category counters, origin weights and growth decay. Source limitations recorded.

08 Colophon

Memory Loss is a digital artwork at chorus.computer, inspired by the National Debt Clock.