1 of 70

AMITY UNIVERSITY KOLKATA · CSIT406 · 3 CREDITS · UG

Social Media

Analytics

From network foundations and data pipelines to machine learning, tools and a hands-on capstone.

Dr. Indraneel Mukhopadhyay

Amity Institute of Information Technology, Amity University Kolkata

Complete classroom & self-study deck · 300 slides · Fully aligned to the CSIT406 syllabus

2 of 70

IV

MODULE

WEIGHTAGE · 25%

Machine Learning in Social Media Analysis

IN THIS MODULE

  • Supervised & unsupervised learning
  • Sentiment classification (TF-IDF, Word2Vec, LSTMs)
  • Fake-news detection & misinformation
  • Recommendation systems & personalisation
  • Social bots & automated-behaviour detection

167

3 of 70

CORE SYLLABUS

MODULE IV · LEARNING OUTCOMES

What You Will Be Able to Do

Apply supervised & unsupervised ML

Choose and train the right model for social-media insight.

Classify sentiment with NLP

Use TF-IDF, Word2Vec and LSTMs for opinion mining.

Detect fake news

Build models that flag misinformation using content and network signals.

Build recommenders

Design content-based and collaborative recommendation systems.

Detect social bots

Identify automated and coordinated inauthentic behaviour.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

168

4 of 70

CORE SYLLABUS

FOUNDATIONS

Why Machine Learning for Social Media?

THE SCALE ARGUMENT

Social platforms produce far too much data for manual analysis. Machine learning finds patterns, classifies content and predicts behaviour automatically — at the scale and speed social media demands.

Volume

Millions of posts per minute — no human can label them.

Speed

Real-time decisions: moderation, ranking, alerts.

Patterns

ML surfaces signals humans would never spot.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

169

5 of 70

DEEPER DIVE

REFRESHER

The Machine-Learning Workflow

1

Define & collect

Frame the task; gather and label data.

2

Preprocess & features

Clean data and engineer features (Module II).

3

Train

Fit a model on the training set.

4

Evaluate

Measure on held-out data; tune.

5

Deploy & monitor

Serve predictions; watch for drift.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

170

6 of 70

CORE SYLLABUS

LEARNING PARADIGMS

Supervised vs. Unsupervised Learning

Supervised

  • Learns from labelled examples
  • Input → known output (label)
  • Tasks: classification, regression
  • e.g. sentiment, fake-news, bot detection
  • Needs a labelled training set

Unsupervised

  • Finds structure without labels
  • Discovers groups and patterns
  • Tasks: clustering, topic modelling
  • e.g. audience segments, themes
  • No labels required

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

171

7 of 70

CORE SYLLABUS

SUPERVISED LEARNING

Supervised Learning

LEARNING FROM LABELS

A supervised model learns a mapping from input features to a known target using labelled examples, then generalises to predict targets for new, unseen inputs.

Features (X)

The inputs — e.g. TF-IDF vectors of a post.

Target (y)

The label — e.g. positive / negative.

Generalisation

Perform well on data it has never seen.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

172

8 of 70

CORE SYLLABUS

SUPERVISED LEARNING

Classification vs. Regression

Classification

Predict a category. Is this tweet positive, negative or neutral? Is this account a bot? Is this news fake?

Regression

Predict a number. How many likes will this post get? What is the expected engagement rate?

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

173

9 of 70

CORE SYLLABUS

SUPERVISED LEARNING

Key Classifiers for Social Data

Naive Bayes

Fast probabilistic baseline; excellent for text.

Logistic Regression

Strong, interpretable linear classifier.

SVM

Powerful margins in high-dimensional text space.

Decision Tree

Interpretable rules; prone to overfit alone.

Random Forest

Robust ensemble of trees; strong default.

Gradient Boosting

XGBoost / LightGBM — top tabular performance.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

174

10 of 70

CORE SYLLABUS

UNSUPERVISED LEARNING

Unsupervised Learning

FINDING HIDDEN STRUCTURE

Unsupervised learning uncovers patterns in unlabelled data — grouping similar users, discovering topics, or reducing dimensions — without being told the right answer.

Clustering

Group similar users or posts.

Topic modelling

Discover latent themes in text.

Dimensionality reduction

Compress features while keeping signal.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

175

11 of 70

CORE SYLLABUS

UNSUPERVISED LEARNING

Clustering Methods

k-Means

Partition into k clusters by nearest centroid. Fast; needs k chosen in advance.

Hierarchical

Build a tree of nested clusters; no k needed, but slower.

DBSCAN

Density-based; finds arbitrary shapes and flags outliers as noise.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

176

12 of 70

DEEPER DIVE

UNSUPERVISED LEARNING

Dimensionality Reduction

FEWER DIMENSIONS, SAME STORY

Text and network features are extremely high-dimensional. Dimensionality reduction compresses them into a few informative dimensions for modelling and visualisation.

PCA

Linear projection onto directions of greatest variance.

t-SNE / UMAP

Non-linear — great for visualising clusters in 2D.

Caution

t-SNE distances between clusters are not meaningful.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

177

13 of 70

DEEPER DIVE

UNSUPERVISED LEARNING

Topic Modelling with LDA

LATENT DIRICHLET ALLOCATION

LDA models each document as a mixture of topics and each topic as a distribution over words, automatically discovering the themes running through a large text corpus.

Topics = word groups

Each topic is a cluster of co-occurring words.

Docs = topic mixes

A post can be 70% politics, 30% sport.

Use

Discover what a community is talking about.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

178

14 of 70

CORE SYLLABUS

MODEL EVALUATION

Train/Test Split & Cross-Validation

HONEST EVALUATION

To estimate real-world performance, we train on one portion of the data and test on a held-out portion the model has never seen. Cross-validation repeats this over several folds for a robust estimate.

Split

Typical 70/30 or 80/20 train/test.

k-fold CV

Rotate the test fold k times, average results.

No leakage

Fit preprocessing on train only, then apply to test.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

179

15 of 70

CORE SYLLABUS

MODEL EVALUATION

Classification Metrics

Metric

What it measures

Use when

Accuracy

Fraction of correct predictions

Classes are balanced

Precision

Of predicted-positive, how many right

False positives costly

Recall

Of actual-positive, how many caught

False negatives costly

F1-score

Harmonic mean of precision & recall

Imbalanced classes

ROC-AUC

Ranking quality across thresholds

Probabilistic output

On social data (bots, fake news) classes are usually imbalanced — prefer F1 and AUC over raw accuracy.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

180

16 of 70

DEEPER DIVE

MODEL EVALUATION

Overfitting, Underfitting & Imbalance

Overfitting

Memorises training noise; great on train, poor on test. Fix: regularise, more data, simpler model.

Underfitting

Too simple to capture the pattern. Fix: richer features, stronger model.

Class imbalance

Rare positives (bots, fake news). Fix: resampling, class weights, F1/AUC.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

181

17 of 70

PRACTICAL

PRACTICAL

A Text-Classification Pipeline

pipeline.py

from sklearn.pipeline import Pipeline

from sklearn.feature_extraction.text \

import TfidfVectorizer

from sklearn.linear_model \

import LogisticRegression

​

clf = Pipeline([

("tfidf", TfidfVectorizer(ngram_range=(1,2))),

("lr", LogisticRegression(max_iter=1000)),

])

clf.fit(X_train, y_train)

print(clf.score(X_test, y_test))

One object

A Pipeline chains vectoriser and classifier — no leakage.

Swap freely

Change LogisticRegression to SVM or NB in one line.

Deployable

The fitted pipeline transforms raw text end-to-end.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

182

18 of 70

CORE SYLLABUS

SENTIMENT CLASSIFICATION

Sentiment Classification as an ML Task

FROM LEXICONS TO LEARNING

Rather than looking up word polarities, we train a classifier on labelled examples so it learns domain-specific sentiment — handling slang, context and negation far better than a fixed lexicon.

Input

Text → numeric features (TF-IDF or embeddings).

Output

Positive / negative / neutral (or a rating).

Advantage

Adapts to your platform and domain.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

183

19 of 70

CORE SYLLABUS

SENTIMENT CLASSIFICATION

Representing Text for Sentiment

THREE LEVELS OF REPRESENTATION

Sentiment models can use sparse counts (Bag of Words), weighted counts (TF-IDF), or dense semantic vectors (word embeddings). Each captures more meaning than the last — at more cost.

BoW

Counts — simple, sparse, order-blind.

TF-IDF

Weighted counts — highlights distinctive words.

Embeddings

Dense meaning — captures similarity and context.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

184

20 of 70

CORE SYLLABUS

SENTIMENT CLASSIFICATION

TF-IDF Features for Sentiment

A STRONG, SIMPLE BASELINE

A TF-IDF vector fed to Logistic Regression or SVM is a remarkably strong sentiment baseline. It is fast, interpretable (you can inspect the most positive/negative words), and hard to beat on small data.

Fast to train

Runs in seconds on thousands of posts.

Interpretable

Model weights reveal sentiment-bearing words.

Limit

No word meaning; “not good” needs bigrams.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

185

21 of 70

CORE SYLLABUS

SENTIMENT CLASSIFICATION · NLP

Word2Vec

WORDS AS DENSE VECTORS

Word2Vec learns a dense vector for every word by predicting context. Two architectures: CBOW predicts a word from its context; Skip-gram predicts context from a word. Similar words end up with similar vectors.

CBOW

Context → target word (fast).

Skip-gram

Word → context (better for rare words).

Result

A geometry where meaning = direction & distance.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

186

22 of 70

WORKED EXAMPLE · EMBEDDING GEOMETRY

vec(“king”) − vec(“man”) + vec(“woman”) ≈ vec(“queen”)

Embeddings place words in a space where analogies become vector arithmetic. “Paris” is to “France” as “Tokyo” is to “Japan” — the same direction. This is why embeddings power modern NLP.

187

23 of 70

DEEPER DIVE

SENTIMENT CLASSIFICATION · NLP

GloVe & FastText

GloVe

Learns embeddings from global word co-occurrence statistics of the whole corpus — complements Word2Vec’s local windows.

FastText

Represents words as bags of character n-grams, so it handles misspellings and rare words — ideal for noisy social text.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

188

24 of 70

CORE SYLLABUS

SENTIMENT CLASSIFICATION · NLP

From Words to Sequences: RNNs

RECURRENT NEURAL NETWORKS

RNNs read text one token at a time, carrying a hidden state that summarises everything seen so far — letting order and context influence the prediction, unlike bag-of-words models.

Sequential

Processes tokens in order, left to right.

Memory

Hidden state carries context forward.

Weakness

Vanishing gradients forget long-range context.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

189

25 of 70

CORE SYLLABUS

SENTIMENT CLASSIFICATION · NLP

LSTMs — Long Short-Term Memory

LSTM

LSTMs are RNNs with gated memory cells that decide what to remember, forget and output. This lets them capture long-range dependencies — crucial for negation and context in sentiment.

Gates

Input, forget and output gates control the cell.

Long context

Remembers “not” many words before the adjective.

Strong for text

A workhorse before transformers took over.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

190

26 of 70

CORE SYLLABUS

SENTIMENT CLASSIFICATION · NLP

An LSTM Sentiment Architecture

1

Embedding

Map each token to a dense vector.

→

2

LSTM layer

Read the sequence, build context.

→

3

Pooling

Summarise into one vector.

→

4

Dense

Fully connected layer.

→

5

Softmax

Output class probabilities.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

191

27 of 70

PRACTICAL

PRACTICAL

A Naive Bayes Sentiment Classifier

nb_sentiment.py

from sklearn.naive_bayes import MultinomialNB

from sklearn.feature_extraction.text \

import TfidfVectorizer

from sklearn.metrics import classification_report

​

vec = TfidfVectorizer(ngram_range=(1,2))

Xtr = vec.fit_transform(train_text)

clf = MultinomialNB().fit(Xtr, y_train)

​

Xte = vec.transform(test_text)

print(classification_report(

y_test, clf.predict(Xte)))

Baseline in minutes

TF-IDF + Naive Bayes is the classic sentiment starting point.

Report, not accuracy

classification_report shows precision, recall and F1 per class.

Then iterate

Beat this baseline before reaching for deep learning.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

192

28 of 70

PRACTICAL

PRACTICAL

An LSTM Sentiment Model (Keras)

lstm_sentiment.py

from tensorflow.keras.models import Sequential

from tensorflow.keras.layers import \

Embedding, LSTM, Dense

​

model = Sequential([

Embedding(vocab, 128),

LSTM(64),

Dense(3, activation="softmax"),

])

model.compile("adam",

"sparse_categorical_crossentropy",

metrics=["accuracy"])

model.fit(X_train, y_train, epochs=5)

Three layers

Embedding → LSTM → softmax captures sequence context.

Needs more data

Deep models beat baselines only with enough labels.

Modern option

For best results today, fine-tune a transformer instead.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

193

29 of 70

DEEPER DIVE

SENTIMENT CLASSIFICATION · NLP

Transformers & BERT

THE MODERN STANDARD

Transformers use self-attention to weigh every word against every other, capturing context in both directions at once. Pre-trained models like BERT are fine-tuned on a small labelled set to reach state-of-the-art sentiment accuracy.

Self-attention

Every token attends to all others.

Pre-trained

Learns language first, your task second.

Fine-tune

A little labelled data goes a long way.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

194

30 of 70

DEEPER DIVE

SENTIMENT CLASSIFICATION

Comparing Sentiment Approaches

Approach

Accuracy

Data need

Cost

Lexicon (VADER)

Low–Med

None

Very low

TF-IDF + LogReg

Medium

Small

Low

Word2Vec + LSTM

Med–High

Large

High

Fine-tuned BERT

High

Medium

High (GPU)

Start simple. A TF-IDF baseline often gets you 80% of the way for a fraction of the effort.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

195

31 of 70

CASE STUDY

SENTIMENT CLASSIFICATION

Case Study — Election Sentiment Analysis

THE SITUATION

Analysts track public sentiment toward candidates during an election by classifying millions of tweets in real time and comparing trends against polling.

Collect & label

Stream tweets by candidate; label a sample to train a domain-specific classifier.

Classify at scale

A fine-tuned model scores sentiment per candidate per hour across regions.

Interpret carefully

Social sentiment ≠ vote share — sampling bias and bots must be corrected for.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

196

32 of 70

CORE SYLLABUS

FAKE NEWS & MISINFORMATION

The Misinformation Problem

FAKE NEWS

Fake news is fabricated or misleading content presented as legitimate news, spread — deliberately or not — through social networks. Detecting it automatically is one of the hardest and most important ML tasks in social media.

Scale & speed

Spreads faster and farther than corrections.

Hard to detect

Style mimics real journalism.

High stakes

Affects health, elections and safety.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

197

33 of 70

CORE SYLLABUS

FAKE NEWS & MISINFORMATION

A Taxonomy of False Information

Fabricated content

Entirely false, invented to deceive.

Misleading content

Real facts framed to mislead.

False context

Genuine content shared with false context.

Imposter content

Impersonates real sources or people.

Satire / parody

No intent to harm, but often mistaken as real.

Manipulated media

Doctored images, video and deepfakes.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

198

34 of 70

CORE SYLLABUS

FAKE NEWS & MISINFORMATION

Four Families of Detection Signals

Content signals

Language, style, sensationalism, clickbait, factual claims and writing quality of the article itself.

Source signals

Credibility and history of the publisher and the account sharing it.

Propagation signals

How it spreads — cascade shape, speed, and who amplifies it.

Social-context signals

User reactions, replies, stance and community response.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

199

35 of 70

CORE SYLLABUS

FAKE NEWS & MISINFORMATION

Content-Based Features

Linguistic cues

Sensational words, excessive punctuation, all-caps, emotional and hyperbolic language.

Readability & style

Grammar quality, complexity and stylistic fingerprints of deceptive writing.

Claim & fact features

Presence of checkable claims, quotes, sources and hedging.

Sentiment & subjectivity

Fake stories skew emotional and highly subjective.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

200

36 of 70

DEEPER DIVE

FAKE NEWS & MISINFORMATION

Propagation-Based Detection

HOW IT SPREADS BETRAYS IT

Even when the text is convincing, the way misinformation propagates differs from real news — different cascade shapes, faster early spread, more bot amplification and distinctive engagement patterns.

Cascade shape

Deep, bursty spread vs. broad, steady sharing.

Amplifiers

Bot and sock-puppet involvement is a strong signal.

Temporal pattern

Unusually fast early velocity flags suspicion.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

201

37 of 70

CORE SYLLABUS

FAKE NEWS & MISINFORMATION

Models for Fake-News Detection

Classical ML

TF-IDF + Logistic Regression / SVM on content features — strong baseline.

Deep learning

LSTMs and CNNs over embeddings capture linguistic patterns.

Transformers

Fine-tuned BERT-style models lead on text-only detection.

Graph models

GNNs over the propagation graph add structural signal.

Multimodal

Combine text, image and network for robustness.

Ensembles

Blend signals to resist adversarial evasion.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

202

38 of 70

PRACTICAL

PRACTICAL

A Fake-News Classifier

fakenews.py

from sklearn.feature_extraction.text \

import TfidfVectorizer

from sklearn.linear_model \

import PassiveAggressiveClassifier

​

vec = TfidfVectorizer(stop_words="english",

max_df=0.7)

Xtr = vec.fit_transform(train_articles)

​

clf = PassiveAggressiveClassifier(max_iter=50)

clf.fit(Xtr, y_train) # labels: REAL / FAKE

Text-only baseline

TF-IDF over article bodies with a linear classifier.

Beware shortcuts

Models can latch onto source style, not truth — test across sources.

Add context

Combine with propagation and source features for robustness.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

203

39 of 70

DEEPER DIVE

FAKE NEWS & MISINFORMATION

Datasets & Benchmarks

LIAR

12.8K human-labelled short political statements with six truth ratings.

FakeNewsNet

Articles plus social context and propagation — for network-aware models.

FEVER

Fact-verification against Wikipedia evidence — claim + evidence.

ISOT / Kaggle

Real vs. fake article corpora for quick baselines.

CoAID / COVID sets

Health misinformation collected during the pandemic.

Caveat

Benchmarks age fast as tactics evolve — validate on fresh data.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

204

40 of 70

DEEPER DIVE

FAKE NEWS & MISINFORMATION

Why Detection Stays Hard

Adversarial evolution

Producers adapt to evade detectors — an arms race.

Deepfakes & synthetic media

AI-generated images, audio and video blur real and fake.

Low-resource languages

Most detectors are English-first; other languages lag.

Truth is contextual

Labelling requires expertise; ground truth is contested and slow.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

205

41 of 70

CASE STUDY

FAKE NEWS & MISINFORMATION

Case Study — The COVID-19 Infodemic

THE SITUATION

During the pandemic, false claims about cures, causes and vaccines spread on social media so fast the WHO called it an “infodemic” — with real consequences for public health.

The spread

Emotional, high-stakes claims cascaded rapidly, often outpacing verified information.

The response

Platforms and researchers deployed classifiers, fact-check labels and friction on sharing.

The lesson

Detection must pair with UX interventions and trusted-source promotion — models alone are not enough.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

206

42 of 70

CORE SYLLABUS

RECOMMENDATION SYSTEMS

Recommendation & Personalisation

RECOMMENDER SYSTEMS

Recommender systems predict what a user will want — which posts to show, accounts to follow, products to buy — personalising the experience and driving the engagement that powers social platforms.

Personalised

Different feed for every user.

Predictive

Estimate the relevance of each item.

High impact

Drives most engagement on modern platforms.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

207

43 of 70

CORE SYLLABUS

RECOMMENDATION SYSTEMS

Content-Based Filtering

RECOMMEND SIMILAR ITEMS

Content-based filtering recommends items similar to those a user liked before, using item features (topics, hashtags, text). It needs no other users — but can trap the user in a narrow bubble.

Item features

Describe each item by its content.

User profile

Aggregate the features they engage with.

Limit

Over-specialisation — little novelty.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

208

44 of 70

CORE SYLLABUS

RECOMMENDATION SYSTEMS

Collaborative Filtering

WISDOM OF SIMILAR USERS

Collaborative filtering recommends items that similar users liked — “people like you also enjoyed…”. User-based finds similar users; item-based finds items co-liked together. It needs no content features at all.

User-based

Find users with similar taste, borrow their likes.

Item-based

Recommend items frequently liked together.

Cold start

Struggles with brand-new users or items.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

209

45 of 70

DEEPER DIVE

RECOMMENDATION SYSTEMS

Matrix Factorisation

LATENT FACTORS

Matrix factorisation decomposes the sparse user–item interaction matrix into low-dimensional user and item vectors, so that their dot product predicts preference — the technique behind the Netflix Prize.

Latent factors

Hidden dimensions of taste and item style.

Predict

Dot product = predicted rating / affinity.

Scales

SVD / ALS handle huge sparse matrices.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

210

46 of 70

DEEPER DIVE

RECOMMENDATION SYSTEMS

Hybrid & Deep Recommenders

Hybrid

Blend content and collaborative signals to cover each other’s weaknesses.

Neural CF

Deep networks learn non-linear user–item interactions.

Graph-based

GNNs over the user–item graph capture higher-order structure.

Sequence models

Model the order of actions to predict the next one.

Context-aware

Use time, place and device to refine recommendations.

Two-tower

Scalable retrieval used in production feeds.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

211

47 of 70

DEEPER DIVE

RECOMMENDATION SYSTEMS

Inside the Social Feed

The feed you see is a ranked recommendation problem solved in milliseconds.

Candidate generation

Retrieve a few thousand plausible items from millions.

Ranking

Score each candidate for predicted engagement.

re-ranking

Apply diversity, freshness and policy rules.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

212

48 of 70

DEEPER DIVE

RECOMMENDATION SYSTEMS

The Hard Problems

Cold start

No history for new users or items — fall back to content or popularity.

Filter bubbles

Over-personalisation narrows exposure and reinforces views.

Diversity vs relevance

The best feed balances what you’ll click with what you should see.

Feedback loops

Recommending what’s popular makes it more popular.

Engagement vs wellbeing

Optimising clicks can harm users — an ethical tension.

Privacy

Personalisation depends on sensitive behavioural data.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

213

49 of 70

PRACTICAL

PRACTICAL

A Simple Collaborative Filter

cf.py

import numpy as np

from sklearn.metrics.pairwise \

import cosine_similarity

​

# R: users x items rating matrix

sim = cosine_similarity(R) # user-user

​

def recommend(u, k=5):

scores = sim[u] @ R

scores[R[u] > 0] = 0 # hide seen

return np.argsort(scores)[::-1][:k]

Similarity first

Cosine similarity finds users with matching taste.

Weighted vote

Similar users’ ratings vote for unseen items.

Then scale

Swap in matrix factorisation (Surprise / implicit) for size.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

214

50 of 70

DEEPER DIVE

RECOMMENDATION SYSTEMS

Evaluating Recommenders

RANKING, NOT JUST RATING

Because users only see a few top items, recommenders are judged on the quality of the ranked top-k list, not just rating accuracy — and ultimately on live engagement via A/B tests.

Precision@k / Recall@k

Relevant items in the top k.

NDCG / MAP

Reward putting the best items highest.

A/B testing

The real test is online user behaviour.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

215

51 of 70

CASE STUDY

RECOMMENDATION SYSTEMS

Case Study — Personalisation at Scale

THE SITUATION

Video and streaming platforms attribute a majority of watch time to recommendations. Their systems choose from billions of items for hundreds of millions of users in real time.

Two stages

Deep candidate generation narrows billions to hundreds; a ranking model orders them per user.

Signals

Watch history, context, freshness and dozens of features feed the ranker.

The tension

Maximising watch time can amplify sensational content — driving investment in responsible-recommendation research.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

216

52 of 70

CORE SYLLABUS

SOCIAL BOTS

What Are Social Bots?

SOCIAL BOT

A social bot is an account controlled wholly or partly by software, posting or interacting automatically. Bots range from helpful (news, weather) to harmful (spam, manipulation, fake amplification).

Automated

Software-driven posting and interaction.

Not all bad

Many bots are useful and declared.

The concern

Deceptive bots distort discourse and metrics.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

217

53 of 70

CORE SYLLABUS

SOCIAL BOTS

Good Bots vs. Bad Bots

Legitimate bots

  • Transparent and declared
  • News, weather and alert bots
  • Customer-service assistants
  • Research and archival bots
  • Add value; follow platform rules

Malicious bots

  • Deceptive — pose as humans
  • Spam and scam distribution
  • Fake followers and engagement
  • Astroturfing and manipulation
  • Amplify misinformation

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

218

54 of 70

CORE SYLLABUS

SOCIAL BOTS

How Bots Distort Discourse

Fake trends

Coordinated posting pushes hashtags to trending.

Amplification

Inflate the apparent popularity of a message.

Fake consensus

Manufacture the illusion of majority opinion.

Skewed sentiment

Distort sentiment analysis and polls.

Misinfo spread

Seed and boost false narratives.

Erode trust

Undermine confidence in online information.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

219

55 of 70

CORE SYLLABUS

SOCIAL BOTS

Bot-Detection Features

Profile features

Account age, follower/following ratio, default avatar, username entropy.

Temporal features

Post frequency, regularity, superhuman activity, timing patterns.

Content features

Repetitive text, template posts, link ratio, low originality.

Network features

Clustering with other bots, coordinated retweets, star patterns.

Engagement features

Unnatural like/retweet ratios and interaction patterns.

Sentiment/style

Uniform tone and copied phrasing across accounts.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

220

56 of 70

CORE SYLLABUS

SOCIAL BOTS

Detection Methods

Supervised

Train on labelled bot/human accounts using the features above — e.g. Random Forest.

Unsupervised

Cluster accounts to surface coordinated, near-identical behaviour.

Botometer

A well-known service scoring the likelihood an account is a bot.

Graph-based

Detect dense, synchronised subgraphs of accounts.

Deep learning

Sequence models over an account’s activity timeline.

Ensembles

Combine signals to resist evasion.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

221

57 of 70

DEEPER DIVE

SOCIAL BOTS

Coordinated Inauthentic Behaviour

BEYOND SINGLE BOTS

The bigger threat is not one bot but networks of accounts acting together — bot farms, sock puppets and troll networks — that coordinate to manipulate. Detecting coordination matters more than flagging individuals.

Coordination

Synchronised posting and identical content.

Timing

Suspiciously simultaneous activity.

Structure

Detect the campaign, not just the accounts.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

222

58 of 70

PRACTICAL

PRACTICAL

Engineering Bot-Likelihood Features

bot_features.py

import pandas as pd

​

def features(acc):

return {

"ff_ratio": acc.following /

max(acc.followers, 1),

"tweets_per_day": acc.tweets /

acc.age_days,

"has_default_pic": acc.default_pic,

"url_ratio": acc.tweets_with_url /

max(acc.tweets, 1),

}

Simple, strong signals

Ratios and rates separate most bots from humans.

Feed a classifier

Assemble features into a table and train Random Forest.

Adversarial

Sophisticated bots mimic humans — combine with network signals.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

223

59 of 70

CASE STUDY

SOCIAL BOTS

Case Study — Bots in an Online Campaign

THE SITUATION

Researchers analysing a politically charged hashtag find a cluster of accounts created on the same day, posting near-identical content in synchronised bursts to inflate a narrative.

The signal

Superhuman posting rates and simultaneous activity across hundreds of accounts.

The structure

A dense retweet subgraph amplifying a small set of seed messages — a coordinated network.

The action

Accounts are flagged and removed; the true organic sentiment is re-estimated without them.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

224

60 of 70

DEEPER DIVE

RESPONSIBLE ML

The Ethics of ML on Social Data

Bias & fairness

Models inherit bias from data — audit across groups.

Transparency

People deserve to know how decisions are made.

Privacy

Inferring traits from posts can violate expectations.

Harm

Moderation errors and profiling cause real harm.

Dual use

The same models detect and enable manipulation.

Accountability

Someone must own the model’s consequences.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

225

61 of 70

DEEPER DIVE

RESPONSIBLE ML

Explainability & Responsible AI

Interpretable models

Prefer transparent models where stakes are high (moderation, credit-like decisions).

Post-hoc explanation

Tools like SHAP and LIME explain individual predictions of complex models.

Fairness audits

Measure error rates across demographic groups; correct disparities.

Human-in-the-loop

Keep people in decisions that affect people.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

226

62 of 70

DEEPER DIVE

BIG PICTURE

The Deep-Learning Landscape for Social Data

CNNs for text

Capture local n-gram patterns efficiently.

RNNs / LSTMs

Model sequence and order in posts.

Transformers

Attention-based; today’s state of the art.

Graph neural nets

Learn over the social graph itself.

Multimodal

Fuse text, image, audio and video.

LLMs

Zero/few-shot classification and generation.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

227

63 of 70

DEEPER DIVE

RESPONSIBLE ML

ML Pitfalls on Social Data — Do & Don’t

Do

  • Start with a simple, strong baseline
  • Use F1/AUC on imbalanced problems
  • Test across sources, time and languages
  • Audit for bias before deploying

Don’t

  • Trust accuracy on rare-class problems
  • Let preprocessing leak test information
  • Assume a model generalises across platforms
  • Deploy a black box for high-stakes calls

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

228

64 of 70

DEEPER DIVE

QUICK REFERENCE

Choosing the Right Model

Task

Start with

Scale up to

Sentiment

TF-IDF + LogReg

Fine-tuned BERT

Fake news

TF-IDF + linear

Multimodal + GNN

Topic discovery

LDA

BERTopic

Recommendation

Collaborative filtering

Neural / two-tower

Bot detection

RF on features

Graph + sequence models

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

229

65 of 70

THE CENTRAL TENSION

The same machine learning that surfaces insight and personalises experience can also manipulate, mislead and discriminate at scale.

Mastering the models is only half the job — wielding them responsibly is the other half.

230

66 of 70

PRACTICAL

PUTTING IT TOGETHER

An End-to-End ML Project Blueprint

1

Frame

Define the task, label scheme and success metric.

2

Data

Collect, clean and label (Modules II–III).

3

Baseline

TF-IDF + linear model; measure with F1/AUC.

4

Improve

Embeddings, deep models, tuning — if warranted.

5

Audit & ship

Check bias, explain, deploy and monitor drift.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

231

67 of 70

DEEPER DIVE

TOOLING

Libraries for ML on Social Data

scikit-learn

Classical ML, pipelines and metrics.

NLTK / spaCy

Text preprocessing and linguistics.

Gensim

Word2Vec, LDA and topic modelling.

TensorFlow / PyTorch

Deep learning for text and graphs.

Hugging Face

Pre-trained transformers, fine-tuning.

Botometer / NDlib

Bot scoring and diffusion simulation.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

232

68 of 70

DEEPER DIVE

QUICK REFERENCE

ML on Social Data — Cheat-Sheet

Concept

One-line reminder

Supervised

Learn from labels; classify or predict a number.

Unsupervised

Find structure — clusters and topics — without labels.

TF-IDF

Weight words by how distinctive they are.

Word2Vec / LSTM

Dense meaning; sequence-aware context.

F1 / AUC

Use these, not accuracy, on imbalanced data.

Recommenders

Content-based, collaborative, or hybrid ranking.

Module IV · Machine Learning in Social Media Analysis

CSIT406 · Social Media Analytics

233

69 of 70

MODULE IV SUMMARY

Machine Learning — Key Takeaways

Two paradigms

Supervised learns from labels; unsupervised finds structure.

Sentiment, from TF-IDF to BERT

Representation drives accuracy; start simple.

Fake news needs many signals

Content, source, propagation and social context.

Recommenders personalise

Content-based, collaborative and hybrid systems.

Bots must be detected

Behavioural and network features expose coordination.

Do it responsibly

Bias, privacy and transparency are part of the job.

234

70 of 70

THANK YOU

Where graph theory meets

business intelligence.

From the structure of networks to the discipline of analytics — you now have the full toolkit to mine social media responsibly and well.

Dr. Indraneel Mukhopadhyay · Amity University Kolkata · CSIT406 Social Media Analytics

300