01 · DATA ANALYSIS & DASHBOARD

Real data.
Prepared with discipline.

Data analysis, cleaning, preparation, dashboarding, and integrity validation for the STRONGR OS recommendation-intelligence workflow.

WHY MIND-SMALL

Public data for real recommendation mechanics.

MIND is a public Microsoft News recommendation dataset based on genuine user-content interaction logs. This capstone uses official MIND-small train/dev data to validate recommendation mechanics without claiming that the model already recommends Strongr Daily content.

Its official chronology supports a real held-out development evaluation rather than a random split.

DATA SCALE

Training candidates5,843,444236,344 positives · 4.044601%
Development candidates2,740,998111,383 positives · 4.063593%
Total candidate exposures8,584,442Official train + development
Prepared Power BI tables
TableRowsPurpose
BehaviorSummary230,117Interaction and history summary
NewsContent65,238Content and metadata context
CategoryEngagementSummary417Category/subcategory engagement view
ContextSummary79Time and candidate-set context
DataQualitySummary15Integrity and readiness checks

DATA PREPARATION

Preserve the source.
Derive the context.

01

Candidate-level records

Original TSVs were preserved; candidate-level rows were created, metadata joined, and historical context derived.

02

Leakage controls

Raw IDs and candidate position were excluded. Train/dev chronology was preserved, with scalers and encoders fit on training only.

03

Explicit missingness

Missing histories and abstracts are represented explicitly rather than silently treated as ordinary content.

DATA QUALITY

0unresolved candidate metadata
0unresolved history metadata
0shared-news metadata conflicts

Key integrity checks: 0 duplicate impression IDs · 0 duplicate news IDs · 0 malformed candidate tokens.

KEY DASHBOARD INSIGHTS

01

Rare positive engagement

Positive engagement is about 4%, making accuracy alone misleading.

02

Stable held-out prevalence

Train and development prevalence are closely aligned.

03

Context matters

Candidate-set size, history, affinity, category/subcategory, and time are important modeling contexts.

04

Preparation passed

Final preparation passed metadata and integrity validation.

POWER BI EVIDENCE

Power BI was used to validate data scale, engagement patterns, recommendation context and model readiness. These presentation cards reproduce the accepted analytical evidence from the validated MIND-small reporting tables.

01 · DATA SCALE

Scale of the prepared recommendation dataset.

02 · ENGAGEMENT

Stable train/dev engagement prevalence.

03 · CANDIDATE CONTEXT

Candidate-set context observed in Power BI.

04 · HISTORY CONTEXT

Interaction-history context observed in Power BI.

05 · CONTENT CONTEXT

Content-category engagement comparison.

06 · DATA INTEGRITY

Validated data-integrity checks.

Original Power BI project retained as the technical source of record.