Skip to main content
andginja

Research notebook · July & August 2026

What the reels data can tell us

Patterns worth investigating, with the limits left in. July’s measured tables and August’s caption-coded model are separate analyses—not a formula for your next reel.

Published Updated

Four counts, different questions

These sample sizes are not interchangeable and must not be added together. August re-coded the existing wider library; it was not a new July table calculation. Commas in the cohort and model tables separate thousands.

Four counts, different questions
VersionReelsWhat was measured
July 20263,527Historical full-metrics sample across 123 hashtags and 21 niches; duration and caption summaries.
July 20261,614Earlier coded subset used for hook, emotion and format summaries; mixed or unknown coding provenance.
August 19, 202611,015Wider library re-coded by Sonnet from caption, source niche and duration; used for model retraining.
August 19, 20263,493Travel subset of the August library; descriptive reach comparisons within travel.

Read associations, not instructions

A median describes the middle observation in a group. It does not forecast your account’s next post. P90 / median compares the 90th percentile with the median: a measure of spread, not statistical variance, a maximum, or the probability of success.

Differences between groups may reflect account size, audience, content, distribution and time since posting. The bars use one neutral colour because a higher median is not proof that changing a feature will improve a reel.

July 2026 · the historical tables

The original figures remain below with their original scope. They have not been recalculated on the 11,015-reel August library. Per-category counts and uncertainty intervals are not available in this published summary.

Duration

July: 3,527 measured reels. Shorter duration groups had higher median views in this sample. This does not establish a universal length or an effect of cutting the same video.
CategoryMedian viewsP90 / median
0-7s40,530
7-15s34,842
15-30s25,994
30-60s22,462
60s+18,240
July: 3,527 measured reels. Shorter duration groups had higher median views in this sample. This does not establish a universal length or an effect of cutting the same video.

Coded emotion

July: earlier 1,614 coded reels. Awe had a large P90 / median ratio (34.7×); this does not mean that most awe-led reels received no views.
CategoryMedian viewsP90 / median
Surprise30,3633.3x
Curiosity27,4176.7x
Satisfaction19,81210.0x
Aspiration17,3106.7x
Amusement16,1679.3x
Awe20,24034.7x
July: earlier 1,614 coded reels. Awe had a large P90 / median ratio (34.7×); this does not mean that most awe-led reels received no views.

Coded hook

July: earlier 1,614 coded reels. These are comparisons between labeled groups, not controlled tests of opening lines.
CategoryMedian viewsP90 / median
Useful list ("5 places that…")34,8425.1x
Disbelief ("this looks unreal")34,4438.4x
Before / after33,34110.2x
Tutorial ("how we do it")32,2945.2x
Immersive POV25,94717.7x
"When you…"18,48921.0x
Question ("which is your fave?")15,5588.5x
No hook13,4827.7x
July: earlier 1,614 coded reels. These are comparisons between labeled groups, not controlled tests of opening lines.

Niche

July historical niche aggregates. Travel had a P90 / median ratio of 40.1×. These selected categories are not representative estimates for every creator in a niche.
CategoryMedian viewsP90 / median
Education / facts31,9404.8x
DIY / craft28,9188.5x
Food & cooking27,3565.5x
Travel25,28440.1x
Comedy22,31916.7x
Satisfying / ASMR12,1218.0x
July historical niche aggregates. Travel had a P90 / median ratio of 40.1×. These selected categories are not representative estimates for every creator in a niche.

Caption length

July: 3,527 measured reels. Historical caption categories are retained; the long category has no precise cutoff in the retained summary. These figures do not isolate a keyword effect.
CategoryMedian viewsP90 / median
Under 40 characters41,043
Long captions (historical category)28,918
100-300 characters23,948
40-100 characters18,646
July: 3,527 measured reels. Historical caption categories are retained; the long category has no precise cutoff in the retained summary. These figures do not isolate a keyword effect.

August 19 · a separate model update

The retrain used 11,015 caption-coded reels and a reach outcome of views ≥ 49,219. Reported test AUC was 0.737, versus train AUC 0.750. The earlier model reported 0.69 on 1,614 coded reels; both sample and labeling changed, so this is not a like-for-like algorithm comparison.

AUC measures ranking discrimination: roughly, the model ranks a randomly selected above-threshold reel higher than a below-threshold reel 73.7% of the time in the reported test sample, with ties receiving half credit. It is not a 73.7% chance that your reel succeeds, an accuracy rate, or the share of views explained by content.

The travel subset contained 3,493 reels, with a reported above-threshold rate of 0.62. The full library’s rate was 0.41. The table reports observed group rates divided by the travel baseline; 1.11× means a group rate 1.11 times that baseline, not an 11% benefit from editing a video.

August travel subset: selected descriptive comparisons (3,493 reels; views ≥ 49,219).
FeatureGroupRate / travel baseline
Duration8–12s1.11×
Duration20–35s0.62×
Duration35s+0.35×
Caption-derived hookDisbelief / awe1.21×
Caption-derived hookImmersive POV1.09×
Caption-derived formatCinematic footage1.03×

Selected categories, not an exhaustive ranking. Category sizes and confidence intervals are not reported here. August’s travel-specific cinematic association does not justify July’s former blanket advice against cinematic footage; cohort, labels and outcome differ.

Method & provenance

July aggregates come from the retained July study. Available records describe 3,527 reels with full metrics and an earlier coded subset of 1,614. The older coder provenance is mixed or unknown; this edition does not claim blinded human rating.

On 2026-08-19, Sonnet assigned a fixed taxonomy using caption text, source niche and duration. These were not labels obtained by watching each video. A logistic regression used L2 regularization and IRLS fitting, one-hot niche/hook/emotion/format features, standardized log duration and log caption length, and an English-language indicator. The reported evaluation used an 80/20 train/test split and rank-based AUC.

Provenance: the internal validation note dated 2026-08-19 and stored model artifact viral-odds-model-2026-08-19.json, interpreted with the 2026-09-05 product-transfer caveat. This September edition publishes a corrected synthesis, not a new model run. Raw scraped records are not published here.

What remains unknown

Sampling. Hashtag-derived observations are not a random or representative sample of Instagram. Account and niche composition, selection and post age can affect comparisons. Posting-day and posting-hour recommendations cannot be established from this summary.

Labels. Caption-derived hook, emotion and format can misdescribe the actual video. Consistent use of one coder does not establish label validity.

Validation. Creator-separated or time-separated holdouts, leakage checks, calibration and uncertainty intervals are not documented in the available summary. Similar training and test AUC does not establish their absence or prove transfer to new creators or future periods.

Outcomes. The model concerns a view threshold, not sends, follows, bookings or revenue. A separate historical exploration of 197 reels from one account is too narrow to turn engagement ratios into general recommendations. No randomized intervention establishes what an edit would cause.

Use this to plan a test

Choose the outcome that matters to your account before choosing a format. For a hotel, that might be qualified enquiries; for a creator, saves or follows. Views alone cannot answer either question.

Use an association as a hypothesis: try a shorter edit when the story allows it, or a clearer opening promise. Compare repeated posts with similar subject matter and record distribution, post age and the business outcome. Keep useful craft judgment separate from measured results; this study does not prescribe eleven seconds, trending audio, or a universal recipe.

Version history

2026-07-30 · First publication. Historical tables from 3,527 measured reels and the earlier 1,614 coded subset.

2026-08-19 · Model artifact. Existing 11,015-reel library re-coded and model retrained; 3,493-reel travel subset reported separately.

2026-09-07 · Article revision. Added the August analysis with explicit cohort boundaries and corrected AUC interpretation. Removed the unvalidated 29%→38% intervention ladder, 31% improvement claim, unsupported blinding and causal advice. Preserved July table values and clarified methods, limitations and the distinction from Video Audit.

Research and interpretation by André Ginja, software engineer and travel content creator.