← Back to Blog
AI & MarketingAIMarketingMarketingAnalyticsContentStrategy

NotebookLM Invented a Chart From Ad Panel Data

I fed six Meta Ads Manager screenshots from an agency account to Google NotebookLM. It returned a twelve-slide deck whose headline claim is a unit error and whose most persuasive chart was drawn from data I never supplied. Here is the audit, and the checklist that came out of it.

SPSantosh Paudel· September 7, 2026· 12 min read· 1 views
Table of contents

It invented the chart because the story needed one. I gave Google NotebookLM six Meta Ads Manager screenshots from an agency ad account I used to run. It returned a twelve-slide deck titled "Anatomy of a Conversion." Slide 8 plots two retention curves falling off a cliff at five seconds. No second-by-second data exists in any of those panels. The deck's headline claim — a "71% discount" — is a unit error. Both arrive in better typography than most human decks I have sat through.

That last part is the actual problem.

What I gave it, and what it gave back

I worked as a creative strategist at a Nepali marketing and production agency, running paid social for large domestic consumer brands — FMCG, retail, that end of the market. The panels come from that account. None names a brand, so no number here is attached to one.

PanelResult metricResultCost per result (as shown)Spend
AReach5,684,111$0.18$1,043.54
BReach569,875$0.05$30.00
CInstagram profile visits111,577$0.02$2,157.09
DInstagram profile visits11,041$0.01$127.87
EInstagram profile visits13,880$0.02$213.82
FReach1,659,478$0.07$114.60

Two caveats. Panel C's spend dwarfs the rest and reads as an account-level aggregate already containing D and E, so the set is not additive and I never sum it. And that second column is the result metric Ads Manager displayed, not the campaign objective, which these screenshots never show. Where I say objective below, I am reading it back off the result metric — an inference, and the same kind this post is about.

NotebookLM picked two for a head-to-head: panel E, named "Campaign Alpha — The Engager", against panel F, "Campaign Beta — The Amplifier". It declared Alpha the winner. Twelve slides, clean hierarchy, a summary that would survive a client meeting.

Error one: the headline is a unit error

The deck's central claim is that "Alpha's sustained attention drove highly qualified traffic at a 71% discount." The 71% is panel E's cost per result set against panel F's, as a ratio: 1 − (0.02 / 0.07) = 0.714.

Those two numbers are not in the same unit; they are separated by a factor of 1,000. Every one of these six panels carries the same "cost per result" label while the denominator underneath it changes. That is my reading of the panels, not a quote from Meta: its help page for the metric renders client-side and I could not pull text out of it, so the arithmetic below is the evidence. Cost per thousand reached is spend over reach, times a thousand — JetMetrics states it as "The amount spent / Reach * 1,000", and DashThis publishes the same shape: "(total advertising cost / number of unique viewers) * 1000". A cost per profile visit is spend over visits, no multiplier.

Doing the division settles which reading applies — what I should have done before reading past the headline.

PanelReadingArithmeticComputes toAds Manager shows
APer thousand reached1,043.54 / (5,684,111 / 1,000)$0.1836$0.18
BPer thousand reached30.00 / (569,875 / 1,000)$0.0526$0.05
FPer thousand reached114.60 / (1,659,478 / 1,000)$0.0691$0.07
CPer single visit2,157.09 / 111,577$0.0193$0.02
DPer single visit127.87 / 11,041$0.0116$0.01
EPer single visit213.82 / 13,880$0.0154$0.02

All six reconcile, and the reach panels only on the per-thousand reading: per single person reached they come to $0.000184, $0.000053 and $0.000069, all of which would display as $0.00.

So panel E's $0.02 is two cents per profile visit; panel F's $0.07 is seven cents per thousand people reached. Dividing one by the other produces a number that means nothing.

Error two: a chart that could not have been drawn

Slide 8 is titled "The critical five-second drop-off cliff." It plots two curves against elapsed seconds; Beta's falls from roughly 74% of viewers to roughly 12% at the five-second mark. A shaded region is labelled "The Engagement Zone: where passive scrollers become active prospects." Those percentages and that label are read off the slide — they are what the chart says, and the chart should not exist. It is the deck's most persuasive slide, and the one that could not have come from my inputs.

The three video panels report four things each.

PanelVideo playsAvg play timeHook rateHold rate
D545,60200:0723%1.12%
E1,001,83000:1028%2.97%
F2,478,11900:0524%1.62%

Average play time is one number, not a curve. A retention curve needs drop-off by second; I supplied a single mean, and no per-second distribution appears anywhere in six screenshots. The curve, the cliff, the shaded zone and the label were produced to illustrate a conclusion, not to display a measurement.

That is the failure mode worth naming. A wrong number is checkable against the source. A sourceless chart is checkable only if you stop and ask which input produced it, and a well-drawn chart discourages the question. Same reason the E-E-A-T problem gets harder when everyone publishes AI content: the surface signals of care no longer indicate care.

Error three: percentage points are not percent

The deck says Alpha's creative was "4% more effective," citing hook rates of 28% and 24%. I have the quote and not the slide number, and I will not invent one — in a post about invented detail, that is the worst error available to me.

The gap between 28% and 24% is four percentage points. As a relative improvement it is 4 / 24 = 16.7%. The deck states the wrong-unit figure as if it were the relative one, understating its own case fourfold. Small on its own, and worth flagging because it is the same mistake a second time: comparing two screenshot numbers without checking they sit on the same scale.

Error four: Beta was judged on a scoreboard it never ran on

The deck concludes Beta underperformed, on cost per profile visit. Panel F's result metric is reach. On unit grounds alone the verdict is invalid.

Correcting the unit does not hand Beta a win. On cost per thousand reached the three reach panels rank B at $0.053, F at $0.069, A at $0.184. Beta is the middle of three, 31% more expensive per thousand reached than panel B — not the standout of the set on its own objective, and not the worst-looking panel either: of the six values on screen, A's $0.18 is the highest. Beta is not the story either way.

One figure does make F look strong, with a warning label. $114.60 bought 2,478,119 video plays — $0.046 per thousand plays, against Alpha's $213.82 for 1,001,830 plays, or $0.213. A 4.6x gap on a consistent basis. But video plays are not F's result metric, and the two panels drove toward different results — which I read as different delivery and different inventory, an inference off these panels rather than anything a platform document told me. Cross-objective and confounded either way: a plays-cost fact, not a verdict.

I cannot tell you Beta's creative was better. I can tell you the deck's basis for saying it was worse is invalid.

What it got right, which is the dangerous part

Where I checked, transcription was accurate: the two cost figures the deck argued from, and the three hook and hold rate pairs, all matched. I did not audit twelve slides number by number, so I am not claiming every figure was right — only that it held everywhere I looked.

The instinct was right too: the retention stage is where these creatives separate. The arithmetic was not.

Definitions first. Vaizle and Exposure both give hook rate as three-second video plays over impressions, and both give hold rate as ThruPlays over three-second plays. These panels cannot be on that second formula: hold rates of 1.12% to 2.97% under hook rates of 23% to 28% only make sense if hold rate here shares the impressions denominator. Inference from the size of the numbers, not a labelled field.

On that reading, hold rate runs through both stages and already contains the hook.

PanelHook rateHold rateRetention given hook (hold / hook)
D23%1.12%4.87%
E28%2.97%10.61%
F24%1.62%6.75%

Hook stage spread: 28 / 23 = 1.22x. Retention given hook: 10.61 / 4.87 = 2.18x. Reported hold rates span 2.65x, which is the two stages multiplied: unrounded, 1.2174 × 2.1783 = 2.65. Round the operands to 1.22 and 2.18 before multiplying and you get 2.66, which is its own small lesson. Quote 2.65x as the retention stage moving and you have counted the hook stage twice. The conclusion survives at the honest figure: the second stage separated these creatives about 1.8x more than the first.

The data also holds a counterexample the deck missed: panel D has a longer average play time than panel F (00:07 vs 00:05) and a lower hold rate (1.12% vs 1.62%). Those two do not move together here, and with n=3 no relationship is established either way.

The design beat most decks I have built by hand, and that is the whole issue. The failure mode is not sloppiness — sloppy work looks sloppy. It is fluency. My read, offered as a read: a well-typeset chart arrives already looking like evidence, and the scepticism it attracts goes to the conclusion rather than the source. I would still use the tool tomorrow. Checked, this deck would have been useful.

The checklist that came out of it

Four questions, twenty minutes, before the deck reaches anyone else.

  1. Reconcile every headline number by hand. Not spot-checks — every number a conclusion rests on. One division killed this headline.
  2. Name the input behind every chart. If you cannot point at the row, cell or screenshot region behind a plotted series, treat the chart as fabricated. Curves are the highest-risk shape: one needs a distribution, and my panels held only aggregates.
  3. Check that compared numbers share a unit. Per-thousand against per-one. Percentage points against percent. Rates with different denominators. Three of the four errors above came from this.
  4. Ask what each number was optimised for. No platform document behind that one; it falls out of the six panels, where one label resolved to two denominators depending on what the campaign chased. Figures produced under different results are confounded; the honest move is to say so.

None of this is specific to NotebookLM. When I scored Claude and ChatGPT against a research and content scorecard, the finding that stuck came from the Tow Center's citation study rather than either model: citation accuracy and answer quality come apart, and a fluent summary tells you nothing about whether its sources exist. A fluent chart tells you nothing about whether its data exists.

The pattern predates AI. My teardown of 389 pages producing 22 clicks on this site — 4,553 impressions, 0.48% CTR, average position 38.8 in the US — was useful less for the analysis than for checking it against the raw export first. The tool just makes the unchecked version arrive faster and looking better.

FAQ

Can NotebookLM analyse data from screenshots accurately?

It transcribed accurately everywhere I checked — the two cost figures the deck argued from, and the three hook and hold rate pairs — which is not the same as auditing all twelve slides. Reasoning about them is where it failed — it compared two costs whose denominators differ by a factor of 1,000, and plotted a chart from data not in the inputs. This is one run of one deck, so I cannot say how often either happens, only that transcription and interpretation failed independently here.

Why did the AI make up a chart?

I cannot see inside the model, so I will describe the output rather than the mechanism. My inputs held one average play time per panel and no per-second drop-off; the deck produced a per-second curve regardless. Nothing in my request forced it to name the input behind each chart, and nothing in its presentation separates the plotted curve from the figures genuinely transcribed.

How do I check an AI-generated marketing report?

Reconcile every headline number by hand, point each chart at the input that produced it, confirm compared figures share a unit, and check whether the campaigns drove toward the same result. If you cannot locate a chart's input, treat the chart as fabricated.

What is the difference between percentage points and percent?

The gap between 24% and 28% is four percentage points. As a relative change it is 16.7%, because four divided by twenty-four is 0.167. Reports saying "4% better" when they mean four points are ambiguous; here it understated the finding fourfold.


Got an AI-generated analysis you are not sure you believe? I will reconcile the headline figures against your source data and tell you which conclusions survive. Quantitative marketing or get in touch.

Free resource

Get the AI Marketing Prompt Pack

30+ tested prompts for images, captions, video scripts, keywords, and full content systems, delivered instantly.

No spam. Unsubscribe anytime.

Browse all free guides →

Want to implement this with guidance?

Santosh helps founders turn insights like this into real systems.

AI Content Systems

External Resources

Further Reading & Tools

Related Posts

01
11 min
SEOAIMarketing
YesterdaySEO

Can Small Brands Win Visibility in AI Answers?

Size correlates with getting cited in AI answers, but the studies cannot show that size causes it. Here is what the citation research actually found, which levers a small brand controls, and which claimed levers failed the only causal test I could find.

Read article
02
11 min
AIAIMarketing
YesterdayAI

Claude vs ChatGPT: A Scorecard for Research and Writing

The honest answer is task-dependent. A weighted rubric for market research, a second one for content creation, verified September 2026 API prices, and the one randomised study I could find on AI-written marketing copy.

Read article
03
8 min
AIAIMarketing
2mo agoAI Marketing

Claude vs ChatGPT for Marketing: 12 Tasks Scored

Twelve named marketing tasks, one rubric, two frontier models, and a token-cost column. The verdict is split, and the method matters more than the scoreboard.

Read article