Vocabulary learning report generated by AI (prompt included)

## Task

Analyze a language learner's flashcard dataset (columns: context, word, meaning, created_at) and produce a monthly vocabulary learning report. created_at marks each card's creation timestamp, functioning as the primary timeline; word-meaning pairs are guaranteed unique, requiring no deduplication check.

## Analysis Instructions

Group rows by month of created_at, merging any month with too few cards into an adjacent month rather than analyzing it as a standalone row.

For each monthly group, infer:

- Theme: dominant topic/domain (e.g. cooking, crafts, nature, daily conversation) → state briefly
- Difficulty: relative difficulty tier (low/mid/high, or a finer gradient if warranted), judged against the overall dataset rather than any external proficiency standard, reflecting the range within that month, not a single point estimate
- Vocabulary type distribution: proportion of nouns, verbs, adjectives, onomatopoeia, idiomatic/set phrases
- Qualitative shift: whether vocabulary within the month leans toward simple standalone words versus nuanced expressions, phrases, or context-dependent usage

## Output Format

Open with a one-line overview: total date range, total card count, any notable gap months.

Follow with a table, one row per month, columns in this exact order:

- period
- difficulty
- card count
- theme, kept short
- 5~10 representative words
- detailed explanation, several sentences minimum

The explanation column MUST cover, in flowing prose: the theme's content/domain characteristics, the reasoning behind the assigned difficulty, the vocabulary type distribution, and any notable change versus the previous month (e.g. a word reappearing with a different meaning, a sudden thematic shift).

## Comprehensive Analysis (below the table)

Write a prose synthesis addressing:

- Difficulty trajectory across the full period: steady increase, spike-then-plateau, or irregular → identify the pattern
- Content/domain consumption pattern: whether thematic focus narrowed, broadened, or shifted over time
- Qualitative vocabulary shift: whether learning progressed from simple word memorization toward idiomatic expressions, nuanced phrases, or contextual usage

Keep this section as connected prose, not a bulleted list; use headers only to separate the table from this synthesis (no deeper subdivision needed).

## Formatting Constraints

Plain text only; no bold or italic styling. Use hyphens for any list, max 2 levels of nesting. Table cells other than the explanation column stay concise (under one sentence); explanation has no such length restriction and must be substantive.

I made a vocabulary learning report using my vocabulary learning history
Just feed your vocabulary learning history data (word, context, meaning, and timestamp) to an LLM you like (GPT/Gemini/Claude) with the prompt above.

You’ll get a table breaking down each month’s theme, difficulty range, representative words, and a comprehensive review for that period. It shows me my learning patterns I didn’t consciously notice and is quite interesting.


I got this data from the Flash card manager of this Add-on. I think you can export your vocabularies from LingQ voca page.

This is the report for my Japanese learning

Monthly Vocabulary Learning Report

The dataset flashcards_ja (1) - flashcards_ja (1).csv covers a total date range from November 22, 2025 to July 5, 2026, containing a total card count of 4,928 cards with no notable gap months.

period difficulty card count theme 5~10 representative words detailed explanation
2025-11 Low range 415 cards Primary daily vocabulary and concrete surroundings. 縞々, 虎, 男性, 麦わら帽子, 撮影, 泳ぐ, 怖い, クラゲ, ヒトデ, 向かう This month establishes the fundamental baseline of the dataset, focusing on beginner-level content. The theme covers immediate concrete surroundings, daily tools, and visible natural elements. The difficulty is designated as low because the cards consist almost entirely of basic vocabulary describing readily observable objects or simple states. In terms of vocabulary type distribution, concrete nouns dominate heavily, alongside elementary active verbs and foundational descriptive adjectives. Since this is the initial period of the dataset, it serves as the benchmark with no previous month for comparison.
2025-12 Low-mid range 733 cards Household objects, body parts, and arcade recreation. 花びら, 足りる, 微妙だ, ヘラ, ギザギザ, 排水管, 首, 腕, 浮き輪, 水筒 The theme expands into specific domestic domains, tracking items found around the house, parts of the human anatomy, and casual amusement activities like arcade games. The difficulty is evaluated as low-mid range because while the vocabulary remains grounded in concrete entities, it incorporates more nuanced spatial and functional items. Nouns continue to represent the largest share of the vocabulary type distribution, though descriptive state-adjectives and functional verbs begin to emerge more frequently. Compared to the previous month, there is a visible transition toward slightly more specialized daily tools and detailed descriptions of environments rather than isolated elementary concepts.
2026-01 Mid range 715 cards Conversational expressions and abstract social terms. 助走, 適当に, 離れる, 旅人, 慣れる, くれぐれも, ずっしりする, 踏む, ちっぽけな, 脅威 The thematic content shifts toward natural conversational interactions, standard colloquial idioms, and descriptions of abstract behavioral traits. The difficulty tier rises to a mid-range level because the vocabulary moves away from visible physical objects into expressions that require contextual comprehension of tone and interpersonal dynamics. The vocabulary type distribution shifts noticeably toward abstract adverbs, state-altering verbs, and set conversational phrases or slang. This marks a sudden qualitative shift from the previous month, moving from concrete household or leisure items to nuanced social speech patterns and figurative language.
2026-02 Mid range 394 cards Conceptual vocabulary, holidays, and cultural traditions. 共通, 推測, 滝, 縁起, 遥かに, 逆に, 込める, 助詞, 伝統, 要素 The content centers on structural language terms, cultural practices, national holidays, and analytical concepts used to describe logic or patterns. The difficulty remains within the mid-range tier, reflecting words required for coherent explanations of cultural rules, societal functions, and comparative reasoning. Abstract nouns and conceptual adverbs comprise the bulk of the vocabulary type distribution, while verbs are primarily used to define relationships between ideas. Compared to the previous month, the focus moves from casual conversational interaction into structured, semi-formal discussions regarding tradition, grammar, and analytical observation.
2026-03 Mid-high range 382 cards Personal monologues, financial actions, and life milestones. 習う, 優れる, 対等に, 極める, まっさらな, 改めて, 雇う, 換金する, 貯める, 屋敷 The theme focuses on individual life management, economic transactions, personal development, and narrative descriptions of experiences abroad or educational methods. The difficulty is classified as mid-high range because it demands familiarity with compound verbs and nuanced descriptors related to mature socio-economic responsibilities. The vocabulary type distribution displays a significant increase in compound action verbs and abstract attributes, while nouns lean toward organizational and environmental settings. This indicates a progression from the previous month’s structured topics into extended, standard-speed speech or monologue streams that handle complex personal and practical scenarios.
2026-04 Mid-high range 433 cards Language learning pedagogy, literacy, and behavioral dynamics. 掛け声, はて, 湧く, まくる, 架空, 儲ける, おごる, たかる, 接する, 目立つ The dominant theme involves the explicit discussion of language acquisition, study strategies, proficiency levels, and social behaviors involving transactions or manipulation. The difficulty is placed in the mid-high tier, as it requires understanding formal educational terminology alongside highly nuanced interpersonal verbs. Nouns representing evaluation criteria and abstract concepts form a major part of the distribution, balanced by dynamic behavioral verbs and occasional rhetorical particles. Compared to the prior period, the vocabulary shows a highly specialized shift toward the meta-commentary of learning itself and the specific vocabulary used in formal Japanese school environments or targeted training.
2026-05 Mid-high to High range 300 cards Professional occupations, modern social habits, and unscripted commentary. 揉む, 機嫌, 気楽, ホッとする, 切り替わる, ほっとく, 整備士, 庭師, 組み込まれる, クビになる The material explores diverse occupational domains, lifestyle choices, emotional management, and commentary on contemporary digital communication habits. The difficulty spans from mid-high to a high range due to the integration of complex administrative vocabulary alongside raw, unscripted colloquialisms. The vocabulary type distribution features a robust blend of technical professional nouns, compound reflexive verbs, and slang or idiomatic set phrases used in casual adult speech. This represents a subtle shift from the academic focus of the previous month into authentic, unfiltered commentary on work-life realities and modern social interactions.
2026-06 ~ 2026-07 High range 1556 cards Cinematic dialogues, screenplays, and dramatic narratives. 上下巻, 綴り, きつい, 大人数, 大豆, イワシ, 太もも, ちぐはぐ, 常連客, たたずむ, 回し者 This extensive period features vocabulary drawn directly from animated feature films and theatrical screenplays, incorporating dramatic, non-standard, and emotionally charged dialogue. The difficulty is assigned to the high tier because the learner is exposed to fast colloquial speech, descriptive literary verbs, archaic expressions, and regional or non-standard variations. The vocabulary type distribution is highly diverse, containing specialized descriptive nouns, heavy sensory onomatopoeia, informal descriptive adjectives, and context-dependent action verbs. This marks a massive and sharp thematic shift from the monologues of previous months, moving directly into full immersion within authentic creative media and rich cinematic narrative structures.

Comprehensive Analysis

The difficulty trajectory across the full period exhibits a steady and progressive increase, moving systematically from elementary foundations to complex, native-level media immersion. In the opening months, the difficulty remains low as the dataset accumulates simple, tangible nouns and basic descriptions. By early 2026, it ascends into a steady mid-range tier where conversational abstractions and structural concepts take precedence, before climbing into mid-high levels that demand comprehension of specialized and professional themes. The trajectory culminates in a high-tier plateau during the final months, driven by an immense volume of advanced cinematic dialogue and theatrical scripts. This continuous upward movement indicates a well-paced advancement in the complexity of the materials being consumed.

The content and domain consumption pattern shows a significant shift over time, expanding from localized immediate environments to broad intellectual monologues, and finally transitioning into authentic narrative fiction. Initially, the thematic focus is narrow and grounded in everyday physical objects, domestic tools, and basic human anatomy. As the months progress, the learner broadens their scope to encompass abstract societal structures, language pedagogy, financial transactions, and professional lifestyles. The final phase represents an immersive expansion into dramatic screenplays, where the domain covers everything from historical fantasy to contemporary slice-of-life cinema. This pattern demonstrates that the learner successfully transitioned from curated, educational domains into unstructured, real-world creative content.

The qualitative vocabulary shift reflects a clear progression from simple standalone word memorization toward highly idiomatic expressions, nuanced phrases, and deeply context-dependent usage. The early entries are predominantly isolated concrete nouns and primary verbs requiring minimal contextual decoding. Throughout the middle months, the vocabulary increasingly incorporates abstract adverbs, compound verbs, and colloquial phrases that alter meaning based on tone and social settings. By the final period, the vocabulary is thoroughly saturated with literary verbs, specific cultural set phrases, sensory onomatopoeia, and dramatic dialogue lines where meaning is entirely intertwined with character context and narrative pacing. This evolution highlights a transition from mechanical vocabulary acquisition to an advanced appreciation of linguistic nuance and natural expression.

3 Likes

I haven’t read the report results, but the idea seems very interesting.

However, how many rows of info the AI would actually handle? Because if you have a few thousand words, okay, but if you have 50 thousand? Are we sure that the LLM will really go through all those lines and do a real analysis? Rather than a sloppy average? Just asking, because the value of the analysis could depend on the type of paid profile the user has. Or other factors.

By the way, this is something useful that LingQ could think about introducing.

Are you able to do a different analysis? More practical?
For example, analysing your database, extracting only the words at level 3, creating a bunch of verified short stories, or even connecting them, using those words, so that you can read/import them with the purpose of strategically converting your level 3 to known. The only problem would be to verify the quality of the stories, but maybe it’s not such a big deal.

Since the topic and difficulty of the words will form clusters in the time series, it seems the data size issue can be resolved through random sampling. In fact, because context is supplementary information that accounts for most of the data size, it appears that even large volumes of data can be processed in a single operation if this is removed.
There are many ways to improve it, but I just made the task most simple.

I did analyze the LingQ words data before; have a look if you are interested.

1 Like

Thanks for the link, I’ll have a look in the next days! :+1: