1 of 18

Beyond the Ballot: TikTok Virality and Political Engagement in Nepal’s 2022 Elections

Master Thesis

Author: Nima Thing

Master of Data Science for Public Policy (2023-2025), Hertie School, Berlin

Supervisor: Prof. Dr. Simon Munzert

Date of Submission: 2025-04-28

2 of 18

Why Study TikTok in Nepal for Elections?

  • Social media has become a key election space for young people in Nepal.
  • TikTok became an important campaign and political attention arena during Nepal’s 2022 local elections.
  • The case is substantively important in political science because some independent candidates gained unusual visibility online.
  • Amidst of long political instability and frustration, TikTok creators mobilized various styles in entertainment platform to support newer voices during election.
  • The case is methodologically important because Nepal is a multilingual, low-resource, under-studied context in political lens (more quantitatively).
  • The broader question shifts from CS “What is TikTok’s hidden algorithm?” CSS - “What observable patterns structure political visibility in Nepal politics?”���

3 of 18

 Research Questions

  • RQ1: How well can political TikTok virality be predicted from observable sender, content, and platform characteristics?
  • RQ2: How are communication styles and political content themes associated with virality?

The thesis separates three tasks:

  • Description of political attention,
  • Prediction from observed metadata,
  • Regression-based associational analysis.

This separation is important because predictive success is not causal explanation.

4 of 18

Data and Insights:

Full collection: 28,165 TikTok videos (Nepal 2022 local elections)

    • Multimodal sample: 2,964 videos
    • Period: March 15 – May 13, 2022

Important caveat:

  • Some sender variables (e.g., total likes, video count) are scrape-time snapshots

5 of 18

 Conceptual Framework

Political virality is treated as platform-mediated political attention.

The framework has three components:

Sender capacity: follower scale, account-level engagement proxies, verification, account age.

Content signaling: communication styles and political content themes.

Platform affordances: timing and platform metadata observable in the dataset.

The framework implies:

  • RQ1 is answered by predictive comparisons across feature groups.
  • RQ2 is answered by associational patterns, not by causal claims about style.

6 of 18

Methods

Virality is operationalized as an engagement-weighted composite score, then used as:

  • a classification target for RQ1,
  • a continuous transformed outcome for RQ2.

RQ1: repeated cross-validated prediction pipeline comparing full, sender-only, content-only, and platform-only models.

RQ2: additive OLS model plus sensitivity checks, interaction model, propensity-score overlap diagnostics, and creator-level variation checks.

Goal:

  • RQ1 identifies what carries predictive information.
  • RQ2 identifies which patterns are associated with higher engagement, while testing whether stronger interpretation is justified.

7 of 18

Exploratory Analysis I

8 of 18

Exploratory Analysis II

9 of 18

RQ1 Main Result

RQ1: Virality is moderately predictable and

sender capacity does most of the work�

  • Full model: AUC 0.800 ± 0.012
  • Sender-only clean: AUC 0.717 ± 0.013
  • Content and platform features still improve

performance by about +0.046 AUC

10 of 18

RQ1 Interpretation

  • Creator-level features (user_like_count, video_count, followers) provide strong

predictive signal for RQ1

  • Content features matter, but less than sender variables
  • What RQ1 does NOT show:
    • TikTok’s ranking algorithm
    • causal effects
    • pure upload-time prediction
  • Important nuance:
    • Middle-virality cases are hardest to classify
  • Takeaway:�Political virality in this dataset is patterned more by observable creator capacity / platform position than by content alone.

sender

platform

11 of 18

RQ2 Main result

RQ2: Actor-Linked Content Themes Show the Clearest Pattern

  • Additive OLS (N = 1,526, R² = 0.127)

Content effects:

    • Independent candidates: β = +0.43 (p < 0.001)
    • Maoist: β = +0.26 (p < 0.05)
    • Congress/UML: not significant
  • Style (pooled model only):
    • Charisma/Leadership: positive
    • Critique/Frustration: positive
    • Others: not significant
    • Interactions do not improve model fit

Actor-linked content (especially independents) is the strongest predictor of higher engagement

12 of 18

Why RQ2 Is Not Causal

Observable overlap is fine: Charisma and non-Charisma videos sit in the same range of creator/timing characteristics (see plot)

  • But overlap alone isn't enough for causal claims
  • Very few creators use more than one style:
    • 1,345 creators total
    • Only 83 use more than one style
    • Only 240 videos provide within-creator variation
  • Within those 83 creators, style effects disappear:
    • Charisma: pooled = positive → within-creator = flips, not significant
    • Critique: pooled = positive → within-creator = loses significance

Conclusion: apparent style effects reflect which creators choose which styles, not style effects themselves

Note*: Small sample, but attenuation pattern is consistent for both pooled-significant styles

13 of 18

What the Thesis is allowing us to infer

  • RQ1 answer: political virality can be predicted moderately well, but mainly from sender capacity.
  • RQ2 answer: higher engagement is more clearly associated with actor-linked political themes, especially independent-candidate content, than with any communication style.

Combined interpretation:

  • On TikTok, political visibility in this dataset is patterned more by who is speaking and which actor they represent than by content form alone.

The thesis therefore supports a framework of:

  • sender capacity first,
  • actor-linked signaling second,
  • platform affordances only indirectly through metadata.

14 of 18

Limitations

Sender features partly retrospective

    • user_like_count and Video_count are scrape-time snapshots, not upload-time values
    • Removing them drops full-model AUC by 0.036 (0.800 → 0.764)
    • Sender-dominance ranking holds either way�

Sample is stratified, not full corpus

    • Multimodal analysis uses 2,964 of 28,165 videos
    • Stratified by view percentile, so representativeness relative to the full corpus not fully validated

RQ2 identifies associations, not causal effects

    • Pre-upload propensity overlap is fine (98.9% common support)
    • But within-creator style variation is sparse: only 83 of 1,345 creators (6.2%) use ≥2 styles, covering 240 videos
    • Style coefficients attenuate 25–40% under creator fixed effects, consistent with creator selection, not causal style effects

15 of 18

Additional Limitations

  • Platform mechanisms remain unobservable
    • TikTok's ranking algorithm, impression counts, and follower/repost graphs are not in the Research API
    • "Platform affordances" in the framework are only approximated through timing and diversification metadata
  • Low-resource NLP ceiling
    • Nepali audio transcript validation Cohen's κ = 0.46 (moderate)
    • Reflects the current limits of multilingual models on Devanagari + Romanized Nepali code-switching, not annotation effort

16 of 18

Future Research

  • Use larger or full-corpus multimodal datasets,
  • Track the same creators longitudinally across campaign periods,
  • Compare TikTok with Facebook, YouTube Shorts, or Instagram Reels,
  • Study network structure and political amplification if better API access becomes available.

17 of 18

Final Takeway

The key issue is not only harmful content, but also unequal political visibility

18 of 18

Q & A