AMITY UNIVERSITY KOLKATA · CSIT406 · 3 CREDITS · UG
Social Media
Analytics
From network foundations and data pipelines to machine learning, tools and a hands-on capstone.
Dr. Indraneel Mukhopadhyay
Amity Institute of Information Technology, Amity University Kolkata
Complete classroom & self-study deck · 300 slides · Fully aligned to the CSIT406 syllabus
IV
MODULE
WEIGHTAGE · 25%
Machine Learning in Social Media Analysis
IN THIS MODULE
167
CORE SYLLABUS
MODULE IV · LEARNING OUTCOMES
What You Will Be Able to Do
Apply supervised & unsupervised ML
Choose and train the right model for social-media insight.
Classify sentiment with NLP
Use TF-IDF, Word2Vec and LSTMs for opinion mining.
Detect fake news
Build models that flag misinformation using content and network signals.
Build recommenders
Design content-based and collaborative recommendation systems.
Detect social bots
Identify automated and coordinated inauthentic behaviour.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
168
CORE SYLLABUS
FOUNDATIONS
Why Machine Learning for Social Media?
THE SCALE ARGUMENT
Social platforms produce far too much data for manual analysis. Machine learning finds patterns, classifies content and predicts behaviour automatically — at the scale and speed social media demands.
Volume
Millions of posts per minute — no human can label them.
Speed
Real-time decisions: moderation, ranking, alerts.
Patterns
ML surfaces signals humans would never spot.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
169
DEEPER DIVE
REFRESHER
The Machine-Learning Workflow
1
Define & collect
Frame the task; gather and label data.
2
Preprocess & features
Clean data and engineer features (Module II).
3
Train
Fit a model on the training set.
4
Evaluate
Measure on held-out data; tune.
5
Deploy & monitor
Serve predictions; watch for drift.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
170
CORE SYLLABUS
LEARNING PARADIGMS
Supervised vs. Unsupervised Learning
Supervised
Unsupervised
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
171
CORE SYLLABUS
SUPERVISED LEARNING
Supervised Learning
LEARNING FROM LABELS
A supervised model learns a mapping from input features to a known target using labelled examples, then generalises to predict targets for new, unseen inputs.
Features (X)
The inputs — e.g. TF-IDF vectors of a post.
Target (y)
The label — e.g. positive / negative.
Generalisation
Perform well on data it has never seen.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
172
CORE SYLLABUS
SUPERVISED LEARNING
Classification vs. Regression
Classification
Predict a category. Is this tweet positive, negative or neutral? Is this account a bot? Is this news fake?
Regression
Predict a number. How many likes will this post get? What is the expected engagement rate?
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
173
CORE SYLLABUS
SUPERVISED LEARNING
Key Classifiers for Social Data
Naive Bayes
Fast probabilistic baseline; excellent for text.
Logistic Regression
Strong, interpretable linear classifier.
SVM
Powerful margins in high-dimensional text space.
Decision Tree
Interpretable rules; prone to overfit alone.
Random Forest
Robust ensemble of trees; strong default.
Gradient Boosting
XGBoost / LightGBM — top tabular performance.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
174
CORE SYLLABUS
UNSUPERVISED LEARNING
Unsupervised Learning
FINDING HIDDEN STRUCTURE
Unsupervised learning uncovers patterns in unlabelled data — grouping similar users, discovering topics, or reducing dimensions — without being told the right answer.
Clustering
Group similar users or posts.
Topic modelling
Discover latent themes in text.
Dimensionality reduction
Compress features while keeping signal.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
175
CORE SYLLABUS
UNSUPERVISED LEARNING
Clustering Methods
k-Means
Partition into k clusters by nearest centroid. Fast; needs k chosen in advance.
Hierarchical
Build a tree of nested clusters; no k needed, but slower.
DBSCAN
Density-based; finds arbitrary shapes and flags outliers as noise.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
176
DEEPER DIVE
UNSUPERVISED LEARNING
Dimensionality Reduction
FEWER DIMENSIONS, SAME STORY
Text and network features are extremely high-dimensional. Dimensionality reduction compresses them into a few informative dimensions for modelling and visualisation.
PCA
Linear projection onto directions of greatest variance.
t-SNE / UMAP
Non-linear — great for visualising clusters in 2D.
Caution
t-SNE distances between clusters are not meaningful.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
177
DEEPER DIVE
UNSUPERVISED LEARNING
Topic Modelling with LDA
LATENT DIRICHLET ALLOCATION
LDA models each document as a mixture of topics and each topic as a distribution over words, automatically discovering the themes running through a large text corpus.
Topics = word groups
Each topic is a cluster of co-occurring words.
Docs = topic mixes
A post can be 70% politics, 30% sport.
Use
Discover what a community is talking about.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
178
CORE SYLLABUS
MODEL EVALUATION
Train/Test Split & Cross-Validation
HONEST EVALUATION
To estimate real-world performance, we train on one portion of the data and test on a held-out portion the model has never seen. Cross-validation repeats this over several folds for a robust estimate.
Split
Typical 70/30 or 80/20 train/test.
k-fold CV
Rotate the test fold k times, average results.
No leakage
Fit preprocessing on train only, then apply to test.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
179
CORE SYLLABUS
MODEL EVALUATION
Classification Metrics
Metric | What it measures | Use when |
Accuracy | Fraction of correct predictions | Classes are balanced |
Precision | Of predicted-positive, how many right | False positives costly |
Recall | Of actual-positive, how many caught | False negatives costly |
F1-score | Harmonic mean of precision & recall | Imbalanced classes |
ROC-AUC | Ranking quality across thresholds | Probabilistic output |
On social data (bots, fake news) classes are usually imbalanced — prefer F1 and AUC over raw accuracy.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
180
DEEPER DIVE
MODEL EVALUATION
Overfitting, Underfitting & Imbalance
Overfitting
Memorises training noise; great on train, poor on test. Fix: regularise, more data, simpler model.
Underfitting
Too simple to capture the pattern. Fix: richer features, stronger model.
Class imbalance
Rare positives (bots, fake news). Fix: resampling, class weights, F1/AUC.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
181
PRACTICAL
PRACTICAL
A Text-Classification Pipeline
pipeline.py
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text \
import TfidfVectorizer
from sklearn.linear_model \
import LogisticRegression
clf = Pipeline([
("tfidf", TfidfVectorizer(ngram_range=(1,2))),
("lr", LogisticRegression(max_iter=1000)),
])
clf.fit(X_train, y_train)
print(clf.score(X_test, y_test))
One object
A Pipeline chains vectoriser and classifier — no leakage.
Swap freely
Change LogisticRegression to SVM or NB in one line.
Deployable
The fitted pipeline transforms raw text end-to-end.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
182
CORE SYLLABUS
SENTIMENT CLASSIFICATION
Sentiment Classification as an ML Task
FROM LEXICONS TO LEARNING
Rather than looking up word polarities, we train a classifier on labelled examples so it learns domain-specific sentiment — handling slang, context and negation far better than a fixed lexicon.
Input
Text → numeric features (TF-IDF or embeddings).
Output
Positive / negative / neutral (or a rating).
Advantage
Adapts to your platform and domain.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
183
CORE SYLLABUS
SENTIMENT CLASSIFICATION
Representing Text for Sentiment
THREE LEVELS OF REPRESENTATION
Sentiment models can use sparse counts (Bag of Words), weighted counts (TF-IDF), or dense semantic vectors (word embeddings). Each captures more meaning than the last — at more cost.
BoW
Counts — simple, sparse, order-blind.
TF-IDF
Weighted counts — highlights distinctive words.
Embeddings
Dense meaning — captures similarity and context.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
184
CORE SYLLABUS
SENTIMENT CLASSIFICATION
TF-IDF Features for Sentiment
A STRONG, SIMPLE BASELINE
A TF-IDF vector fed to Logistic Regression or SVM is a remarkably strong sentiment baseline. It is fast, interpretable (you can inspect the most positive/negative words), and hard to beat on small data.
Fast to train
Runs in seconds on thousands of posts.
Interpretable
Model weights reveal sentiment-bearing words.
Limit
No word meaning; “not good” needs bigrams.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
185
CORE SYLLABUS
SENTIMENT CLASSIFICATION · NLP
Word2Vec
WORDS AS DENSE VECTORS
Word2Vec learns a dense vector for every word by predicting context. Two architectures: CBOW predicts a word from its context; Skip-gram predicts context from a word. Similar words end up with similar vectors.
CBOW
Context → target word (fast).
Skip-gram
Word → context (better for rare words).
Result
A geometry where meaning = direction & distance.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
186
WORKED EXAMPLE · EMBEDDING GEOMETRY
vec(“king”) − vec(“man”) + vec(“woman”) ≈ vec(“queen”)
Embeddings place words in a space where analogies become vector arithmetic. “Paris” is to “France” as “Tokyo” is to “Japan” — the same direction. This is why embeddings power modern NLP.
187
DEEPER DIVE
SENTIMENT CLASSIFICATION · NLP
GloVe & FastText
GloVe
Learns embeddings from global word co-occurrence statistics of the whole corpus — complements Word2Vec’s local windows.
FastText
Represents words as bags of character n-grams, so it handles misspellings and rare words — ideal for noisy social text.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
188
CORE SYLLABUS
SENTIMENT CLASSIFICATION · NLP
From Words to Sequences: RNNs
RECURRENT NEURAL NETWORKS
RNNs read text one token at a time, carrying a hidden state that summarises everything seen so far — letting order and context influence the prediction, unlike bag-of-words models.
Sequential
Processes tokens in order, left to right.
Memory
Hidden state carries context forward.
Weakness
Vanishing gradients forget long-range context.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
189
CORE SYLLABUS
SENTIMENT CLASSIFICATION · NLP
LSTMs — Long Short-Term Memory
LSTM
LSTMs are RNNs with gated memory cells that decide what to remember, forget and output. This lets them capture long-range dependencies — crucial for negation and context in sentiment.
Gates
Input, forget and output gates control the cell.
Long context
Remembers “not” many words before the adjective.
Strong for text
A workhorse before transformers took over.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
190
CORE SYLLABUS
SENTIMENT CLASSIFICATION · NLP
An LSTM Sentiment Architecture
1
Embedding
Map each token to a dense vector.
→
2
LSTM layer
Read the sequence, build context.
→
3
Pooling
Summarise into one vector.
→
4
Dense
Fully connected layer.
→
5
Softmax
Output class probabilities.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
191
PRACTICAL
PRACTICAL
A Naive Bayes Sentiment Classifier
nb_sentiment.py
from sklearn.naive_bayes import MultinomialNB
from sklearn.feature_extraction.text \
import TfidfVectorizer
from sklearn.metrics import classification_report
vec = TfidfVectorizer(ngram_range=(1,2))
Xtr = vec.fit_transform(train_text)
clf = MultinomialNB().fit(Xtr, y_train)
Xte = vec.transform(test_text)
print(classification_report(
y_test, clf.predict(Xte)))
Baseline in minutes
TF-IDF + Naive Bayes is the classic sentiment starting point.
Report, not accuracy
classification_report shows precision, recall and F1 per class.
Then iterate
Beat this baseline before reaching for deep learning.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
192
PRACTICAL
PRACTICAL
An LSTM Sentiment Model (Keras)
lstm_sentiment.py
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import \
Embedding, LSTM, Dense
model = Sequential([
Embedding(vocab, 128),
LSTM(64),
Dense(3, activation="softmax"),
])
model.compile("adam",
"sparse_categorical_crossentropy",
metrics=["accuracy"])
model.fit(X_train, y_train, epochs=5)
Three layers
Embedding → LSTM → softmax captures sequence context.
Needs more data
Deep models beat baselines only with enough labels.
Modern option
For best results today, fine-tune a transformer instead.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
193
DEEPER DIVE
SENTIMENT CLASSIFICATION · NLP
Transformers & BERT
THE MODERN STANDARD
Transformers use self-attention to weigh every word against every other, capturing context in both directions at once. Pre-trained models like BERT are fine-tuned on a small labelled set to reach state-of-the-art sentiment accuracy.
Self-attention
Every token attends to all others.
Pre-trained
Learns language first, your task second.
Fine-tune
A little labelled data goes a long way.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
194
DEEPER DIVE
SENTIMENT CLASSIFICATION
Comparing Sentiment Approaches
Approach | Accuracy | Data need | Cost |
Lexicon (VADER) | Low–Med | None | Very low |
TF-IDF + LogReg | Medium | Small | Low |
Word2Vec + LSTM | Med–High | Large | High |
Fine-tuned BERT | High | Medium | High (GPU) |
Start simple. A TF-IDF baseline often gets you 80% of the way for a fraction of the effort.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
195
CASE STUDY
SENTIMENT CLASSIFICATION
Case Study — Election Sentiment Analysis
THE SITUATION
Analysts track public sentiment toward candidates during an election by classifying millions of tweets in real time and comparing trends against polling.
Collect & label
Stream tweets by candidate; label a sample to train a domain-specific classifier.
Classify at scale
A fine-tuned model scores sentiment per candidate per hour across regions.
Interpret carefully
Social sentiment ≠ vote share — sampling bias and bots must be corrected for.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
196
CORE SYLLABUS
FAKE NEWS & MISINFORMATION
The Misinformation Problem
FAKE NEWS
Fake news is fabricated or misleading content presented as legitimate news, spread — deliberately or not — through social networks. Detecting it automatically is one of the hardest and most important ML tasks in social media.
Scale & speed
Spreads faster and farther than corrections.
Hard to detect
Style mimics real journalism.
High stakes
Affects health, elections and safety.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
197
CORE SYLLABUS
FAKE NEWS & MISINFORMATION
A Taxonomy of False Information
Fabricated content
Entirely false, invented to deceive.
Misleading content
Real facts framed to mislead.
False context
Genuine content shared with false context.
Imposter content
Impersonates real sources or people.
Satire / parody
No intent to harm, but often mistaken as real.
Manipulated media
Doctored images, video and deepfakes.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
198
CORE SYLLABUS
FAKE NEWS & MISINFORMATION
Four Families of Detection Signals
Content signals
Language, style, sensationalism, clickbait, factual claims and writing quality of the article itself.
Source signals
Credibility and history of the publisher and the account sharing it.
Propagation signals
How it spreads — cascade shape, speed, and who amplifies it.
Social-context signals
User reactions, replies, stance and community response.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
199
CORE SYLLABUS
FAKE NEWS & MISINFORMATION
Content-Based Features
Linguistic cues
Sensational words, excessive punctuation, all-caps, emotional and hyperbolic language.
Readability & style
Grammar quality, complexity and stylistic fingerprints of deceptive writing.
Claim & fact features
Presence of checkable claims, quotes, sources and hedging.
Sentiment & subjectivity
Fake stories skew emotional and highly subjective.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
200
DEEPER DIVE
FAKE NEWS & MISINFORMATION
Propagation-Based Detection
HOW IT SPREADS BETRAYS IT
Even when the text is convincing, the way misinformation propagates differs from real news — different cascade shapes, faster early spread, more bot amplification and distinctive engagement patterns.
Cascade shape
Deep, bursty spread vs. broad, steady sharing.
Amplifiers
Bot and sock-puppet involvement is a strong signal.
Temporal pattern
Unusually fast early velocity flags suspicion.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
201
CORE SYLLABUS
FAKE NEWS & MISINFORMATION
Models for Fake-News Detection
Classical ML
TF-IDF + Logistic Regression / SVM on content features — strong baseline.
Deep learning
LSTMs and CNNs over embeddings capture linguistic patterns.
Transformers
Fine-tuned BERT-style models lead on text-only detection.
Graph models
GNNs over the propagation graph add structural signal.
Multimodal
Combine text, image and network for robustness.
Ensembles
Blend signals to resist adversarial evasion.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
202
PRACTICAL
PRACTICAL
A Fake-News Classifier
fakenews.py
from sklearn.feature_extraction.text \
import TfidfVectorizer
from sklearn.linear_model \
import PassiveAggressiveClassifier
vec = TfidfVectorizer(stop_words="english",
max_df=0.7)
Xtr = vec.fit_transform(train_articles)
clf = PassiveAggressiveClassifier(max_iter=50)
clf.fit(Xtr, y_train) # labels: REAL / FAKE
Text-only baseline
TF-IDF over article bodies with a linear classifier.
Beware shortcuts
Models can latch onto source style, not truth — test across sources.
Add context
Combine with propagation and source features for robustness.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
203
DEEPER DIVE
FAKE NEWS & MISINFORMATION
Datasets & Benchmarks
LIAR
12.8K human-labelled short political statements with six truth ratings.
FakeNewsNet
Articles plus social context and propagation — for network-aware models.
FEVER
Fact-verification against Wikipedia evidence — claim + evidence.
ISOT / Kaggle
Real vs. fake article corpora for quick baselines.
CoAID / COVID sets
Health misinformation collected during the pandemic.
Caveat
Benchmarks age fast as tactics evolve — validate on fresh data.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
204
DEEPER DIVE
FAKE NEWS & MISINFORMATION
Why Detection Stays Hard
Adversarial evolution
Producers adapt to evade detectors — an arms race.
Deepfakes & synthetic media
AI-generated images, audio and video blur real and fake.
Low-resource languages
Most detectors are English-first; other languages lag.
Truth is contextual
Labelling requires expertise; ground truth is contested and slow.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
205
CASE STUDY
FAKE NEWS & MISINFORMATION
Case Study — The COVID-19 Infodemic
THE SITUATION
During the pandemic, false claims about cures, causes and vaccines spread on social media so fast the WHO called it an “infodemic” — with real consequences for public health.
The spread
Emotional, high-stakes claims cascaded rapidly, often outpacing verified information.
The response
Platforms and researchers deployed classifiers, fact-check labels and friction on sharing.
The lesson
Detection must pair with UX interventions and trusted-source promotion — models alone are not enough.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
206
CORE SYLLABUS
RECOMMENDATION SYSTEMS
Recommendation & Personalisation
RECOMMENDER SYSTEMS
Recommender systems predict what a user will want — which posts to show, accounts to follow, products to buy — personalising the experience and driving the engagement that powers social platforms.
Personalised
Different feed for every user.
Predictive
Estimate the relevance of each item.
High impact
Drives most engagement on modern platforms.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
207
CORE SYLLABUS
RECOMMENDATION SYSTEMS
Content-Based Filtering
RECOMMEND SIMILAR ITEMS
Content-based filtering recommends items similar to those a user liked before, using item features (topics, hashtags, text). It needs no other users — but can trap the user in a narrow bubble.
Item features
Describe each item by its content.
User profile
Aggregate the features they engage with.
Limit
Over-specialisation — little novelty.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
208
CORE SYLLABUS
RECOMMENDATION SYSTEMS
Collaborative Filtering
WISDOM OF SIMILAR USERS
Collaborative filtering recommends items that similar users liked — “people like you also enjoyed…”. User-based finds similar users; item-based finds items co-liked together. It needs no content features at all.
User-based
Find users with similar taste, borrow their likes.
Item-based
Recommend items frequently liked together.
Cold start
Struggles with brand-new users or items.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
209
DEEPER DIVE
RECOMMENDATION SYSTEMS
Matrix Factorisation
LATENT FACTORS
Matrix factorisation decomposes the sparse user–item interaction matrix into low-dimensional user and item vectors, so that their dot product predicts preference — the technique behind the Netflix Prize.
Latent factors
Hidden dimensions of taste and item style.
Predict
Dot product = predicted rating / affinity.
Scales
SVD / ALS handle huge sparse matrices.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
210
DEEPER DIVE
RECOMMENDATION SYSTEMS
Hybrid & Deep Recommenders
Hybrid
Blend content and collaborative signals to cover each other’s weaknesses.
Neural CF
Deep networks learn non-linear user–item interactions.
Graph-based
GNNs over the user–item graph capture higher-order structure.
Sequence models
Model the order of actions to predict the next one.
Context-aware
Use time, place and device to refine recommendations.
Two-tower
Scalable retrieval used in production feeds.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
211
DEEPER DIVE
RECOMMENDATION SYSTEMS
Inside the Social Feed
The feed you see is a ranked recommendation problem solved in milliseconds.
Candidate generation
Retrieve a few thousand plausible items from millions.
Ranking
Score each candidate for predicted engagement.
re-ranking
Apply diversity, freshness and policy rules.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
212
DEEPER DIVE
RECOMMENDATION SYSTEMS
The Hard Problems
Cold start
No history for new users or items — fall back to content or popularity.
Filter bubbles
Over-personalisation narrows exposure and reinforces views.
Diversity vs relevance
The best feed balances what you’ll click with what you should see.
Feedback loops
Recommending what’s popular makes it more popular.
Engagement vs wellbeing
Optimising clicks can harm users — an ethical tension.
Privacy
Personalisation depends on sensitive behavioural data.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
213
PRACTICAL
PRACTICAL
A Simple Collaborative Filter
cf.py
import numpy as np
from sklearn.metrics.pairwise \
import cosine_similarity
# R: users x items rating matrix
sim = cosine_similarity(R) # user-user
def recommend(u, k=5):
scores = sim[u] @ R
scores[R[u] > 0] = 0 # hide seen
return np.argsort(scores)[::-1][:k]
Similarity first
Cosine similarity finds users with matching taste.
Weighted vote
Similar users’ ratings vote for unseen items.
Then scale
Swap in matrix factorisation (Surprise / implicit) for size.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
214
DEEPER DIVE
RECOMMENDATION SYSTEMS
Evaluating Recommenders
RANKING, NOT JUST RATING
Because users only see a few top items, recommenders are judged on the quality of the ranked top-k list, not just rating accuracy — and ultimately on live engagement via A/B tests.
Precision@k / Recall@k
Relevant items in the top k.
NDCG / MAP
Reward putting the best items highest.
A/B testing
The real test is online user behaviour.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
215
CASE STUDY
RECOMMENDATION SYSTEMS
Case Study — Personalisation at Scale
THE SITUATION
Video and streaming platforms attribute a majority of watch time to recommendations. Their systems choose from billions of items for hundreds of millions of users in real time.
Two stages
Deep candidate generation narrows billions to hundreds; a ranking model orders them per user.
Signals
Watch history, context, freshness and dozens of features feed the ranker.
The tension
Maximising watch time can amplify sensational content — driving investment in responsible-recommendation research.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
216
CORE SYLLABUS
SOCIAL BOTS
What Are Social Bots?
SOCIAL BOT
A social bot is an account controlled wholly or partly by software, posting or interacting automatically. Bots range from helpful (news, weather) to harmful (spam, manipulation, fake amplification).
Automated
Software-driven posting and interaction.
Not all bad
Many bots are useful and declared.
The concern
Deceptive bots distort discourse and metrics.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
217
CORE SYLLABUS
SOCIAL BOTS
Good Bots vs. Bad Bots
Legitimate bots
Malicious bots
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
218
CORE SYLLABUS
SOCIAL BOTS
How Bots Distort Discourse
Fake trends
Coordinated posting pushes hashtags to trending.
Amplification
Inflate the apparent popularity of a message.
Fake consensus
Manufacture the illusion of majority opinion.
Skewed sentiment
Distort sentiment analysis and polls.
Misinfo spread
Seed and boost false narratives.
Erode trust
Undermine confidence in online information.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
219
CORE SYLLABUS
SOCIAL BOTS
Bot-Detection Features
Profile features
Account age, follower/following ratio, default avatar, username entropy.
Temporal features
Post frequency, regularity, superhuman activity, timing patterns.
Content features
Repetitive text, template posts, link ratio, low originality.
Network features
Clustering with other bots, coordinated retweets, star patterns.
Engagement features
Unnatural like/retweet ratios and interaction patterns.
Sentiment/style
Uniform tone and copied phrasing across accounts.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
220
CORE SYLLABUS
SOCIAL BOTS
Detection Methods
Supervised
Train on labelled bot/human accounts using the features above — e.g. Random Forest.
Unsupervised
Cluster accounts to surface coordinated, near-identical behaviour.
Botometer
A well-known service scoring the likelihood an account is a bot.
Graph-based
Detect dense, synchronised subgraphs of accounts.
Deep learning
Sequence models over an account’s activity timeline.
Ensembles
Combine signals to resist evasion.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
221
DEEPER DIVE
SOCIAL BOTS
Coordinated Inauthentic Behaviour
BEYOND SINGLE BOTS
The bigger threat is not one bot but networks of accounts acting together — bot farms, sock puppets and troll networks — that coordinate to manipulate. Detecting coordination matters more than flagging individuals.
Coordination
Synchronised posting and identical content.
Timing
Suspiciously simultaneous activity.
Structure
Detect the campaign, not just the accounts.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
222
PRACTICAL
PRACTICAL
Engineering Bot-Likelihood Features
bot_features.py
import pandas as pd
def features(acc):
return {
"ff_ratio": acc.following /
max(acc.followers, 1),
"tweets_per_day": acc.tweets /
acc.age_days,
"has_default_pic": acc.default_pic,
"url_ratio": acc.tweets_with_url /
max(acc.tweets, 1),
}
Simple, strong signals
Ratios and rates separate most bots from humans.
Feed a classifier
Assemble features into a table and train Random Forest.
Adversarial
Sophisticated bots mimic humans — combine with network signals.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
223
CASE STUDY
SOCIAL BOTS
Case Study — Bots in an Online Campaign
THE SITUATION
Researchers analysing a politically charged hashtag find a cluster of accounts created on the same day, posting near-identical content in synchronised bursts to inflate a narrative.
The signal
Superhuman posting rates and simultaneous activity across hundreds of accounts.
The structure
A dense retweet subgraph amplifying a small set of seed messages — a coordinated network.
The action
Accounts are flagged and removed; the true organic sentiment is re-estimated without them.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
224
DEEPER DIVE
RESPONSIBLE ML
The Ethics of ML on Social Data
Bias & fairness
Models inherit bias from data — audit across groups.
Transparency
People deserve to know how decisions are made.
Privacy
Inferring traits from posts can violate expectations.
Harm
Moderation errors and profiling cause real harm.
Dual use
The same models detect and enable manipulation.
Accountability
Someone must own the model’s consequences.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
225
DEEPER DIVE
RESPONSIBLE ML
Explainability & Responsible AI
Interpretable models
Prefer transparent models where stakes are high (moderation, credit-like decisions).
Post-hoc explanation
Tools like SHAP and LIME explain individual predictions of complex models.
Fairness audits
Measure error rates across demographic groups; correct disparities.
Human-in-the-loop
Keep people in decisions that affect people.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
226
DEEPER DIVE
BIG PICTURE
The Deep-Learning Landscape for Social Data
CNNs for text
Capture local n-gram patterns efficiently.
RNNs / LSTMs
Model sequence and order in posts.
Transformers
Attention-based; today’s state of the art.
Graph neural nets
Learn over the social graph itself.
Multimodal
Fuse text, image, audio and video.
LLMs
Zero/few-shot classification and generation.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
227
DEEPER DIVE
RESPONSIBLE ML
ML Pitfalls on Social Data — Do & Don’t
Do
Don’t
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
228
DEEPER DIVE
QUICK REFERENCE
Choosing the Right Model
Task | Start with | Scale up to |
Sentiment | TF-IDF + LogReg | Fine-tuned BERT |
Fake news | TF-IDF + linear | Multimodal + GNN |
Topic discovery | LDA | BERTopic |
Recommendation | Collaborative filtering | Neural / two-tower |
Bot detection | RF on features | Graph + sequence models |
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
229
THE CENTRAL TENSION
The same machine learning that surfaces insight and personalises experience can also manipulate, mislead and discriminate at scale.
Mastering the models is only half the job — wielding them responsibly is the other half.
230
PRACTICAL
PUTTING IT TOGETHER
An End-to-End ML Project Blueprint
1
Frame
Define the task, label scheme and success metric.
2
Data
Collect, clean and label (Modules II–III).
3
Baseline
TF-IDF + linear model; measure with F1/AUC.
4
Improve
Embeddings, deep models, tuning — if warranted.
5
Audit & ship
Check bias, explain, deploy and monitor drift.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
231
DEEPER DIVE
TOOLING
Libraries for ML on Social Data
scikit-learn
Classical ML, pipelines and metrics.
NLTK / spaCy
Text preprocessing and linguistics.
Gensim
Word2Vec, LDA and topic modelling.
TensorFlow / PyTorch
Deep learning for text and graphs.
Hugging Face
Pre-trained transformers, fine-tuning.
Botometer / NDlib
Bot scoring and diffusion simulation.
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
232
DEEPER DIVE
QUICK REFERENCE
ML on Social Data — Cheat-Sheet
Concept | One-line reminder |
Supervised | Learn from labels; classify or predict a number. |
Unsupervised | Find structure — clusters and topics — without labels. |
TF-IDF | Weight words by how distinctive they are. |
Word2Vec / LSTM | Dense meaning; sequence-aware context. |
F1 / AUC | Use these, not accuracy, on imbalanced data. |
Recommenders | Content-based, collaborative, or hybrid ranking. |
Module IV · Machine Learning in Social Media Analysis
CSIT406 · Social Media Analytics
233
MODULE IV SUMMARY
Machine Learning — Key Takeaways
Two paradigms
Supervised learns from labels; unsupervised finds structure.
Sentiment, from TF-IDF to BERT
Representation drives accuracy; start simple.
Fake news needs many signals
Content, source, propagation and social context.
Recommenders personalise
Content-based, collaborative and hybrid systems.
Bots must be detected
Behavioural and network features expose coordination.
Do it responsibly
Bias, privacy and transparency are part of the job.
234
THANK YOU
Where graph theory meets
business intelligence.
From the structure of networks to the discipline of analytics — you now have the full toolkit to mine social media responsibly and well.
Dr. Indraneel Mukhopadhyay · Amity University Kolkata · CSIT406 Social Media Analytics
300