NotebookLM Invented a Chart From Ad Panel Data
I fed six Meta Ads Manager screenshots from an agency account to Google NotebookLM. It returned a twelve-slide deck whose headline claim is a unit error and whose most persuasive chart was drawn from data I never supplied. Here is the audit, and the checklist that came out of it.
Table of contents
It invented the chart because the story needed one. I gave Google NotebookLM six Meta Ads Manager screenshots from an agency ad account I used to run. It returned a twelve-slide deck titled "Anatomy of a Conversion." Slide 8 plots two retention curves falling off a cliff at five seconds. No second-by-second data exists in any of those panels. The deck's headline claim — a "71% discount" — is a unit error. Both arrive in better typography than most human decks I have sat through.
That last part is the actual problem.
What I gave it, and what it gave back
I worked as a creative strategist at a Nepali marketing and production agency, running paid social for large domestic consumer brands — FMCG, retail, that end of the market. The panels come from that account. None names a brand, so no number here is attached to one.
| Panel | Result metric | Result | Cost per result (as shown) | Spend |
|---|---|---|---|---|
| A | Reach | 5,684,111 | $0.18 | $1,043.54 |
| B | Reach | 569,875 | $0.05 | $30.00 |
| C | Instagram profile visits | 111,577 | $0.02 | $2,157.09 |
| D | Instagram profile visits | 11,041 | $0.01 | $127.87 |
| E | Instagram profile visits | 13,880 | $0.02 | $213.82 |
| F | Reach | 1,659,478 | $0.07 | $114.60 |
Two caveats. Panel C's spend dwarfs the rest and reads as an account-level aggregate already containing D and E, so the set is not additive and I never sum it. And that second column is the result metric Ads Manager displayed, not the campaign objective, which these screenshots never show. Where I say objective below, I am reading it back off the result metric — an inference, and the same kind this post is about.
NotebookLM picked two for a head-to-head: panel E, named "Campaign Alpha — The Engager", against panel F, "Campaign Beta — The Amplifier". It declared Alpha the winner. Twelve slides, clean hierarchy, a summary that would survive a client meeting.
Error one: the headline is a unit error
The deck's central claim is that "Alpha's sustained attention drove highly qualified traffic at a 71% discount." The 71% is panel E's cost per result set against panel F's, as a ratio: 1 − (0.02 / 0.07) = 0.714.
Those two numbers are not in the same unit; they are separated by a factor of 1,000. Every one of these six panels carries the same "cost per result" label while the denominator underneath it changes. That is my reading of the panels, not a quote from Meta: its help page for the metric renders client-side and I could not pull text out of it, so the arithmetic below is the evidence. Cost per thousand reached is spend over reach, times a thousand — JetMetrics states it as "The amount spent / Reach * 1,000", and DashThis publishes the same shape: "(total advertising cost / number of unique viewers) * 1000". A cost per profile visit is spend over visits, no multiplier.
Doing the division settles which reading applies — what I should have done before reading past the headline.
| Panel | Reading | Arithmetic | Computes to | Ads Manager shows |
|---|---|---|---|---|
| A | Per thousand reached | 1,043.54 / (5,684,111 / 1,000) | $0.1836 | $0.18 |
| B | Per thousand reached | 30.00 / (569,875 / 1,000) | $0.0526 | $0.05 |
| F | Per thousand reached | 114.60 / (1,659,478 / 1,000) | $0.0691 | $0.07 |
| C | Per single visit | 2,157.09 / 111,577 | $0.0193 | $0.02 |
| D | Per single visit | 127.87 / 11,041 | $0.0116 | $0.01 |
| E | Per single visit | 213.82 / 13,880 | $0.0154 | $0.02 |
All six reconcile, and the reach panels only on the per-thousand reading: per single person reached they come to $0.000184, $0.000053 and $0.000069, all of which would display as $0.00.
So panel E's $0.02 is two cents per profile visit; panel F's $0.07 is seven cents per thousand people reached. Dividing one by the other produces a number that means nothing.
Error two: a chart that could not have been drawn
Slide 8 is titled "The critical five-second drop-off cliff." It plots two curves against elapsed seconds; Beta's falls from roughly 74% of viewers to roughly 12% at the five-second mark. A shaded region is labelled "The Engagement Zone: where passive scrollers become active prospects." Those percentages and that label are read off the slide — they are what the chart says, and the chart should not exist. It is the deck's most persuasive slide, and the one that could not have come from my inputs.
The three video panels report four things each.
| Panel | Video plays | Avg play time | Hook rate | Hold rate |
|---|---|---|---|---|
| D | 545,602 | 00:07 | 23% | 1.12% |
| E | 1,001,830 | 00:10 | 28% | 2.97% |
| F | 2,478,119 | 00:05 | 24% | 1.62% |
Average play time is one number, not a curve. A retention curve needs drop-off by second; I supplied a single mean, and no per-second distribution appears anywhere in six screenshots. The curve, the cliff, the shaded zone and the label were produced to illustrate a conclusion, not to display a measurement.
That is the failure mode worth naming. A wrong number is checkable against the source. A sourceless chart is checkable only if you stop and ask which input produced it, and a well-drawn chart discourages the question. Same reason the E-E-A-T problem gets harder when everyone publishes AI content: the surface signals of care no longer indicate care.
Error three: percentage points are not percent
The deck says Alpha's creative was "4% more effective," citing hook rates of 28% and 24%. I have the quote and not the slide number, and I will not invent one — in a post about invented detail, that is the worst error available to me.
The gap between 28% and 24% is four percentage points. As a relative improvement it is 4 / 24 = 16.7%. The deck states the wrong-unit figure as if it were the relative one, understating its own case fourfold. Small on its own, and worth flagging because it is the same mistake a second time: comparing two screenshot numbers without checking they sit on the same scale.
Error four: Beta was judged on a scoreboard it never ran on
The deck concludes Beta underperformed, on cost per profile visit. Panel F's result metric is reach. On unit grounds alone the verdict is invalid.
Correcting the unit does not hand Beta a win. On cost per thousand reached the three reach panels rank B at $0.053, F at $0.069, A at $0.184. Beta is the middle of three, 31% more expensive per thousand reached than panel B — not the standout of the set on its own objective, and not the worst-looking panel either: of the six values on screen, A's $0.18 is the highest. Beta is not the story either way.
One figure does make F look strong, with a warning label. $114.60 bought 2,478,119 video plays — $0.046 per thousand plays, against Alpha's $213.82 for 1,001,830 plays, or $0.213. A 4.6x gap on a consistent basis. But video plays are not F's result metric, and the two panels drove toward different results — which I read as different delivery and different inventory, an inference off these panels rather than anything a platform document told me. Cross-objective and confounded either way: a plays-cost fact, not a verdict.
I cannot tell you Beta's creative was better. I can tell you the deck's basis for saying it was worse is invalid.
What it got right, which is the dangerous part
Where I checked, transcription was accurate: the two cost figures the deck argued from, and the three hook and hold rate pairs, all matched. I did not audit twelve slides number by number, so I am not claiming every figure was right — only that it held everywhere I looked.
The instinct was right too: the retention stage is where these creatives separate. The arithmetic was not.
Definitions first. Vaizle and Exposure both give hook rate as three-second video plays over impressions, and both give hold rate as ThruPlays over three-second plays. These panels cannot be on that second formula: hold rates of 1.12% to 2.97% under hook rates of 23% to 28% only make sense if hold rate here shares the impressions denominator. Inference from the size of the numbers, not a labelled field.
On that reading, hold rate runs through both stages and already contains the hook.
| Panel | Hook rate | Hold rate | Retention given hook (hold / hook) |
|---|---|---|---|
| D | 23% | 1.12% | 4.87% |
| E | 28% | 2.97% | 10.61% |
| F | 24% | 1.62% | 6.75% |
Hook stage spread: 28 / 23 = 1.22x. Retention given hook: 10.61 / 4.87 = 2.18x. Reported hold rates span 2.65x, which is the two stages multiplied: unrounded, 1.2174 × 2.1783 = 2.65. Round the operands to 1.22 and 2.18 before multiplying and you get 2.66, which is its own small lesson. Quote 2.65x as the retention stage moving and you have counted the hook stage twice. The conclusion survives at the honest figure: the second stage separated these creatives about 1.8x more than the first.
The data also holds a counterexample the deck missed: panel D has a longer average play time than panel F (00:07 vs 00:05) and a lower hold rate (1.12% vs 1.62%). Those two do not move together here, and with n=3 no relationship is established either way.
The design beat most decks I have built by hand, and that is the whole issue. The failure mode is not sloppiness — sloppy work looks sloppy. It is fluency. My read, offered as a read: a well-typeset chart arrives already looking like evidence, and the scepticism it attracts goes to the conclusion rather than the source. I would still use the tool tomorrow. Checked, this deck would have been useful.
The checklist that came out of it
Four questions, twenty minutes, before the deck reaches anyone else.
- —Reconcile every headline number by hand. Not spot-checks — every number a conclusion rests on. One division killed this headline.
- —Name the input behind every chart. If you cannot point at the row, cell or screenshot region behind a plotted series, treat the chart as fabricated. Curves are the highest-risk shape: one needs a distribution, and my panels held only aggregates.
- —Check that compared numbers share a unit. Per-thousand against per-one. Percentage points against percent. Rates with different denominators. Three of the four errors above came from this.
- —Ask what each number was optimised for. No platform document behind that one; it falls out of the six panels, where one label resolved to two denominators depending on what the campaign chased. Figures produced under different results are confounded; the honest move is to say so.
None of this is specific to NotebookLM. When I scored Claude and ChatGPT against a research and content scorecard, the finding that stuck came from the Tow Center's citation study rather than either model: citation accuracy and answer quality come apart, and a fluent summary tells you nothing about whether its sources exist. A fluent chart tells you nothing about whether its data exists.
The pattern predates AI. My teardown of 389 pages producing 22 clicks on this site — 4,553 impressions, 0.48% CTR, average position 38.8 in the US — was useful less for the analysis than for checking it against the raw export first. The tool just makes the unchecked version arrive faster and looking better.
FAQ
Can NotebookLM analyse data from screenshots accurately?
It transcribed accurately everywhere I checked — the two cost figures the deck argued from, and the three hook and hold rate pairs — which is not the same as auditing all twelve slides. Reasoning about them is where it failed — it compared two costs whose denominators differ by a factor of 1,000, and plotted a chart from data not in the inputs. This is one run of one deck, so I cannot say how often either happens, only that transcription and interpretation failed independently here.
Why did the AI make up a chart?
I cannot see inside the model, so I will describe the output rather than the mechanism. My inputs held one average play time per panel and no per-second drop-off; the deck produced a per-second curve regardless. Nothing in my request forced it to name the input behind each chart, and nothing in its presentation separates the plotted curve from the figures genuinely transcribed.
How do I check an AI-generated marketing report?
Reconcile every headline number by hand, point each chart at the input that produced it, confirm compared figures share a unit, and check whether the campaigns drove toward the same result. If you cannot locate a chart's input, treat the chart as fabricated.
What is the difference between percentage points and percent?
The gap between 24% and 28% is four percentage points. As a relative change it is 16.7%, because four divided by twenty-four is 0.167. Reports saying "4% better" when they mean four points are ambiguous; here it understated the finding fourfold.
Got an AI-generated analysis you are not sure you believe? I will reconcile the headline figures against your source data and tell you which conclusions survive. Quantitative marketing or get in touch.
Get the AI Marketing Prompt Pack
30+ tested prompts for images, captions, video scripts, keywords, and full content systems, delivered instantly.
Browse all free guides →Want to implement this with guidance?
Santosh helps founders turn insights like this into real systems.
External Resources
Further Reading & Tools
Content Marketing Institute
Annual content marketing benchmarks — budgets, channels, and outcomes
HubSpot Marketing Blog
Data-driven marketing research, inbound strategy, and content guides
Semrush Blog
SEO, content, and digital marketing research and strategy guides
Marketing Week
UK's leading marketing news, strategy insight, and industry research