Research notebook · July & August 2026
What the reels data can tell us
Patterns worth investigating, with the limits left in. July’s measured tables and August’s caption-coded model are separate analyses—not a formula for your next reel.
Published Updated
Four counts, different questions
These sample sizes are not interchangeable and must not be added together. August re-coded the existing wider library; it was not a new July table calculation. Commas in the cohort and model tables separate thousands.
| Version | Reels | What was measured |
|---|---|---|
| July 2026 | 3,527 | Historical full-metrics sample across 123 hashtags and 21 niches; duration and caption summaries. |
| July 2026 | 1,614 | Earlier coded subset used for hook, emotion and format summaries; mixed or unknown coding provenance. |
| August 19, 2026 | 11,015 | Wider library re-coded by Sonnet from caption, source niche and duration; used for model retraining. |
| August 19, 2026 | 3,493 | Travel subset of the August library; descriptive reach comparisons within travel. |
Read associations, not instructions
A median describes the middle observation in a group. It does not forecast your account’s next post. P90 / median compares the 90th percentile with the median: a measure of spread, not statistical variance, a maximum, or the probability of success.
Differences between groups may reflect account size, audience, content, distribution and time since posting. The bars use one neutral colour because a higher median is not proof that changing a feature will improve a reel.
July 2026 · the historical tables
The original figures remain below with their original scope. They have not been recalculated on the 11,015-reel August library. Per-category counts and uncertainty intervals are not available in this published summary.
Duration
| Category | Median views | P90 / median |
|---|---|---|
| 0-7s | 40,530 | – |
| 7-15s | 34,842 | – |
| 15-30s | 25,994 | – |
| 30-60s | 22,462 | – |
| 60s+ | 18,240 | – |
Coded emotion
| Category | Median views | P90 / median |
|---|---|---|
| Surprise | 30,363 | 3.3x |
| Curiosity | 27,417 | 6.7x |
| Satisfaction | 19,812 | 10.0x |
| Aspiration | 17,310 | 6.7x |
| Amusement | 16,167 | 9.3x |
| Awe | 20,240 | 34.7x |
Coded hook
| Category | Median views | P90 / median |
|---|---|---|
| Useful list ("5 places that…") | 34,842 | 5.1x |
| Disbelief ("this looks unreal") | 34,443 | 8.4x |
| Before / after | 33,341 | 10.2x |
| Tutorial ("how we do it") | 32,294 | 5.2x |
| Immersive POV | 25,947 | 17.7x |
| "When you…" | 18,489 | 21.0x |
| Question ("which is your fave?") | 15,558 | 8.5x |
| No hook | 13,482 | 7.7x |
Niche
| Category | Median views | P90 / median |
|---|---|---|
| Education / facts | 31,940 | 4.8x |
| DIY / craft | 28,918 | 8.5x |
| Food & cooking | 27,356 | 5.5x |
| Travel | 25,284 | 40.1x |
| Comedy | 22,319 | 16.7x |
| Satisfying / ASMR | 12,121 | 8.0x |
Caption length
| Category | Median views | P90 / median |
|---|---|---|
| Under 40 characters | 41,043 | – |
| Long captions (historical category) | 28,918 | – |
| 100-300 characters | 23,948 | – |
| 40-100 characters | 18,646 | – |
August 19 · a separate model update
The retrain used 11,015 caption-coded reels and a reach outcome of views ≥ 49,219. Reported test AUC was 0.737, versus train AUC 0.750. The earlier model reported 0.69 on 1,614 coded reels; both sample and labeling changed, so this is not a like-for-like algorithm comparison.
AUC measures ranking discrimination: roughly, the model ranks a randomly selected above-threshold reel higher than a below-threshold reel 73.7% of the time in the reported test sample, with ties receiving half credit. It is not a 73.7% chance that your reel succeeds, an accuracy rate, or the share of views explained by content.
The travel subset contained 3,493 reels, with a reported above-threshold rate of 0.62. The full library’s rate was 0.41. The table reports observed group rates divided by the travel baseline; 1.11× means a group rate 1.11 times that baseline, not an 11% benefit from editing a video.
| Feature | Group | Rate / travel baseline |
|---|---|---|
| Duration | 8–12s | 1.11× |
| Duration | 20–35s | 0.62× |
| Duration | 35s+ | 0.35× |
| Caption-derived hook | Disbelief / awe | 1.21× |
| Caption-derived hook | Immersive POV | 1.09× |
| Caption-derived format | Cinematic footage | 1.03× |
Selected categories, not an exhaustive ranking. Category sizes and confidence intervals are not reported here. August’s travel-specific cinematic association does not justify July’s former blanket advice against cinematic footage; cohort, labels and outcome differ.
Method & provenance
July aggregates come from the retained July study. Available records describe 3,527 reels with full metrics and an earlier coded subset of 1,614. The older coder provenance is mixed or unknown; this edition does not claim blinded human rating.
On 2026-08-19, Sonnet assigned a fixed taxonomy using caption text, source niche and duration. These were not labels obtained by watching each video. A logistic regression used L2 regularization and IRLS fitting, one-hot niche/hook/emotion/format features, standardized log duration and log caption length, and an English-language indicator. The reported evaluation used an 80/20 train/test split and rank-based AUC.
Provenance: the internal validation note dated 2026-08-19 and stored model artifact viral-odds-model-2026-08-19.json, interpreted with the 2026-09-05 product-transfer caveat. This September edition publishes a corrected synthesis, not a new model run. Raw scraped records are not published here.
What remains unknown
Sampling. Hashtag-derived observations are not a random or representative sample of Instagram. Account and niche composition, selection and post age can affect comparisons. Posting-day and posting-hour recommendations cannot be established from this summary.
Labels. Caption-derived hook, emotion and format can misdescribe the actual video. Consistent use of one coder does not establish label validity.
Validation. Creator-separated or time-separated holdouts, leakage checks, calibration and uncertainty intervals are not documented in the available summary. Similar training and test AUC does not establish their absence or prove transfer to new creators or future periods.
Outcomes. The model concerns a view threshold, not sends, follows, bookings or revenue. A separate historical exploration of 197 reels from one account is too narrow to turn engagement ratios into general recommendations. No randomized intervention establishes what an edit would cause.
Use this to plan a test
Choose the outcome that matters to your account before choosing a format. For a hotel, that might be qualified enquiries; for a creator, saves or follows. Views alone cannot answer either question.
Use an association as a hypothesis: try a shorter edit when the story allows it, or a clearer opening promise. Compare repeated posts with similar subject matter and record distribution, post age and the business outcome. Keep useful craft judgment separate from measured results; this study does not prescribe eleven seconds, trending audio, or a universal recipe.
Version history
2026-07-30 · First publication. Historical tables from 3,527 measured reels and the earlier 1,614 coded subset.
2026-08-19 · Model artifact. Existing 11,015-reel library re-coded and model retrained; 3,493-reel travel subset reported separately.
2026-09-07 · Article revision. Added the August analysis with explicit cohort boundaries and corrected AUC interpretation. Removed the unvalidated 29%→38% intervention ladder, 31% improvement claim, unsupported blinding and causal advice. Preserved July table values and clarified methods, limitations and the distinction from Video Audit.