1 of 56

AMITY UNIVERSITY KOLKATA · CSIT406 · 3 CREDITS · UG

Social Media

Analytics

From network foundations and data pipelines to machine learning, tools and a hands-on capstone.

Dr. Indraneel Mukhopadhyay

Amity Institute of Information Technology, Amity University Kolkata

Complete classroom & self-study deck · 300 slides · Fully aligned to the CSIT406 syllabus

2 of 56

V

MODULE

WEIGHTAGE · 20%

Tools & Practical Implementation

IN THIS MODULE

  • NetworkX, Gephi, Tweepy, Scrapy
  • Sentiment Analysis APIs
  • Hands-on projects with real datasets
  • Capstone: analysing & visualising trends
  • Building an end-to-end pipeline in Python

235

3 of 56

CORE SYLLABUS

MODULE V · LEARNING OUTCOMES

What You Will Be Able to Do

Use NetworkX & Gephi

Build, analyse and visualise networks programmatically and interactively.

Collect data with Tweepy & Scrapy

Gather real social-media data through APIs and scraping.

Apply sentiment APIs

Integrate ready-made sentiment services and local models.

Run hands-on projects

Complete mini-projects on real-world datasets.

Deliver a capstone

Analyse and visualise social-media trends end-to-end in Python.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

236

4 of 56

CORE SYLLABUS

THE TOOLKIT

The Practitioner’s Toolkit

Five core tools from the syllabus, each owning a stage of the analytics pipeline.

Scrapy

Scrape web data at scale.

Tweepy

Collect data via the X API.

Sentiment APIs

Score opinion & emotion.

NetworkX

Analyse graphs in Python.

Gephi

Visualise networks.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

237

5 of 56

DEEPER DIVE

THE TOOLKIT

The Python Data Ecosystem

pandas / NumPy

Load, clean and manipulate tabular data.

NetworkX / igraph

Graph construction and algorithms.

matplotlib / seaborn

Static charts and statistical plots.

NLTK / spaCy

Text preprocessing and NLP.

scikit-learn

Machine-learning models and metrics.

Plotly / Gephi

Interactive and network visualisation.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

238

6 of 56

CORE SYLLABUS

NETWORKX

NetworkX

GRAPHS IN PYTHON

NetworkX is the standard Python library for creating, manipulating and analysing graphs. It offers rich data structures for networks and dozens of built-in algorithms — the analytical engine of this course.

Flexible graphs

Directed, undirected, weighted, multigraphs.

Algorithms

Centrality, paths, communities, link prediction.

Integrates

Plays well with pandas, matplotlib and Gephi.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

239

7 of 56

PRACTICAL

NETWORKX

Core Operations

nx_basics.py

import networkx as nx

G = nx.DiGraph()

G.add_node("A", country="IN")

G.add_edge("A", "B", weight=3)

G.add_edges_from([("B","C"),("C","A")])

print(G.nodes(data=True))

print(G.in_degree("A"))

print(list(G.successors("A")))

print(nx.shortest_path(G, "A", "C"))

Nodes carry data

Attach attributes to nodes and edges for richer analysis.

Directed methods

in_degree, successors and predecessors respect direction.

Algorithms built in

shortest_path and hundreds more are one call away.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

240

8 of 56

PRACTICAL

NETWORKX

Metrics & Algorithms

nx_metrics.py

import networkx as nx

nx.density(G)

nx.degree_centrality(G)

nx.betweenness_centrality(G)

nx.pagerank(G, weight="weight")

nx.average_clustering(G)

nx.connected_components(G.to_undirected())

from networkx.algorithms.community \

import greedy_modularity_communities

greedy_modularity_communities(G)

Whole course in calls

Every metric from Modules I–III is a NetworkX function.

Ranking made easy

Sort the returned dicts to shortlist influencers.

Communities too

Built-in modularity communities for quick detection.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

241

9 of 56

PRACTICAL

NETWORKX

Quick Visualisation

nx_draw.py

import matplotlib.pyplot as plt

import networkx as nx

pos = nx.spring_layout(G, seed=42)

deg = dict(G.degree())

nx.draw_networkx(

G, pos, with_labels=True,

node_size=[300*deg[n] for n in G],

node_color="#7C5CFC", edge_color="#ccc")

plt.axis("off"); plt.show()

Layout first

spring_layout positions nodes by a force simulation.

Encode metrics

Size nodes by degree to reveal hubs at a glance.

For big graphs

Export to Gephi — matplotlib struggles past a few hundred nodes.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

242

10 of 56

DEEPER DIVE

NETWORKX

Strengths & Limits

Strengths

  • Pure-Python, easy to learn
  • Huge algorithm library
  • Great for analysis and scripting
  • Integrates with the data stack

Limits

  • Slow on very large graphs (millions of edges)
  • Weak built-in visualisation
  • For scale, consider igraph, graph-tool or Spark
  • Not a database — hold graphs in memory

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

243

11 of 56

CORE SYLLABUS

GEPHI

Gephi

INTERACTIVE NETWORK VISUALISATION

Gephi is a free, open-source desktop application for exploring and visualising networks interactively. Where NetworkX computes, Gephi reveals — turning large graphs into readable, beautiful maps.

Visual-first

Explore structure by sight, not just numbers.

Built-in stats

Modularity, centrality and more, no code.

Interoperable

Imports GEXF/GraphML exported from NetworkX.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

244

12 of 56

CORE SYLLABUS

GEPHI

The Gephi Workflow

1

Import

Load a GEXF/CSV edge list.

2

Layout

Run ForceAtlas2 to spread nodes.

3

Stats

Compute modularity & centrality.

4

Style

Colour by community, size by degree.

5

Export

Publish a high-res figure.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

245

13 of 56

DEEPER DIVE

GEPHI

Layout Algorithms

ForceAtlas2

Force-directed layout that pulls connected nodes together and pushes others apart — the Gephi default for revealing communities.

Fruchterman–Reingold

Classic force-directed layout producing balanced, aesthetically even graphs.

Yifan Hu / OpenOrd

Scalable layouts for large graphs that emphasise clustering.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

246

14 of 56

PRACTICAL

GEPHI

Styling & Interpreting Visuals

Size = importance

Map node size to degree or betweenness so influencers pop out.

Colour = community

Colour nodes by modularity class to make groups visible.

Filter for clarity

Hide low-degree nodes and giant components to reduce clutter.

Read the map

Dense clumps are communities; central nodes are brokers; isolates sit at the edges.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

247

15 of 56

DEEPER DIVE

TOOLING

NetworkX vs. Gephi — When to Use Which

Dimension

NetworkX

Gephi

Interface

Code (Python)

Interactive GUI

Best at

Computation & automation

Exploration & visuals

Reproducible

Yes (scripts)

Less so (manual)

Scale

Medium

Large (with layouts)

Workflow

Compute → export

Import → visualise

Common pattern: compute in NetworkX, export GEXF, then visualise and refine in Gephi.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

248

16 of 56

CORE SYLLABUS

TWEEPY

Tweepy

THE PYTHON X/TWITTER CLIENT

Tweepy is a Python library that wraps the X (Twitter) API, handling authentication, requests, pagination and rate limits so you can collect tweets, users and relationships with clean, readable code.

Handles auth

Manages OAuth and bearer tokens for you.

Search & stream

Query recent tweets or stream in real time.

Rate-limit aware

Can wait and resume automatically.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

249

17 of 56

PRACTICAL

TWEEPY

Authentication & Setup

tweepy_setup.py

import tweepy, os

# keep secrets in environment variables

client = tweepy.Client(

bearer_token=os.environ["X_BEARER"],

wait_on_rate_limit=True)

me = client.get_user(username="AmityUni")

print(me.data.id, me.data.name)

Never hard-code keys

Load credentials from environment variables, not source.

Auto back-off

wait_on_rate_limit pauses instead of crashing.

Respect terms

Collect only what the developer agreement allows.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

250

18 of 56

PRACTICAL

TWEEPY

Collecting & Paginating Data

tweepy_collect.py

import tweepy, pandas as pd

rows = []

for tw in tweepy.Paginator(

client.search_recent_tweets,

query="#DataScience lang:en -is:retweet",

tweet_fields=["created_at","public_metrics"],

max_results=100).flatten(limit=1000):

rows.append([tw.id, tw.text,

tw.created_at])

pd.DataFrame(rows).to_csv("tweets.csv")

Paginator handles pages

flatten() gives a simple loop over many pages.

Save as you go

Persist to CSV/Parquet for reproducible analysis.

Filter at source

Query operators cut noise before it reaches you.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

251

19 of 56

PRACTICAL

TWEEPY

Building a Reusable Dataset

Define the query

Fix hashtags, language, date range and exclusions up front.

Store raw + clean

Keep a raw copy; write a separate cleaned table for analysis.

Extract edges

Derive mention/retweet edges to build the graph for Module III.

Log provenance

Record when and how you collected — essential for reproducibility.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

252

20 of 56

CORE SYLLABUS

SCRAPY

Scrapy

A WEB-SCRAPING FRAMEWORK

Scrapy is a fast, extensible Python framework for large-scale web crawling. It manages requests, concurrency, retries and data pipelines — the tool of choice when no API exists and scraping is permitted.

Fast & async

Concurrent requests out of the box.

Pipelines

Clean, validate and store scraped items.

Extensible

Middleware for proxies, throttling and more.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

253

21 of 56

DEEPER DIVE

SCRAPY

Scrapy Architecture

Spiders

Define which URLs to crawl and how to parse them into items.

Items

Structured containers for the fields you extract.

Item pipelines

Process items — clean, dedupe, validate, store.

Middlewares

Hook into requests/responses for headers, proxies, retries.

Scheduler

Queues and prioritises requests efficiently.

Settings

Throttle rate, obey robots.txt, set user agents.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

254

22 of 56

PRACTICAL

SCRAPY

A Simple Spider

quotes_spider.py

import scrapy

class PostSpider(scrapy.Spider):

name = "posts"

start_urls = ["https://example.com/feed"]

def parse(self, response):

for p in response.css(".post"):

yield {

"title": p.css("h2::text").get(),

"likes": p.css(".likes::text").get(),

}

CSS/XPath selectors

Pinpoint the fields to extract from each page.

yield items

Yielded dicts flow into pipelines and output files.

Run it

scrapy crawl posts -o posts.json — done.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

255

23 of 56

DEEPER DIVE

SCRAPY

Scraping Responsibly

Check robots.txt & ToS

Respect what a site permits; some prohibit scraping entirely.

Throttle politely

Set download delays and concurrency limits to avoid overloading servers.

Mind personal data

Scraping personal data triggers DPDP/GDPR obligations — anonymise.

Prefer APIs

If an API exists, use it — it is cleaner, safer and usually legal.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

256

24 of 56

CORE SYLLABUS

SENTIMENT APIS

Sentiment Analysis APIs

SENTIMENT AS A SERVICE

Sentiment APIs let you send text and receive polarity or emotion scores without training a model — trading control and cost for speed and convenience. They complement the local models of Module IV.

No training

Production-grade results instantly.

Multilingual

Many languages supported out of the box.

Trade-offs

Cost, privacy and less domain control.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

257

25 of 56

DEEPER DIVE

SENTIMENT APIS

Popular Sentiment Services

Service

Type

Notes

Google Cloud NL

Cloud API

Sentiment, entities, syntax; multilingual

AWS Comprehend

Cloud API

Sentiment, key phrases, PII detection

Azure Language

Cloud API

Sentiment + opinion mining (aspect-level)

Hugging Face

Hosted models

Thousands of open models; self-host option

VADER / TextBlob

Local library

Free, offline, great for teaching

For sensitive data, prefer local models (VADER, Hugging Face) so text never leaves your environment.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

258

26 of 56

CORE SYLLABUS

SENTIMENT APIS

Local, Free Sentiment Tools

VADER

Rule-based and tuned for social media — handles emojis, slang and emphasis. Returns a compound score in [−1, +1]. No training, no cost, fully offline.

TextBlob

Beginner-friendly API returning polarity and subjectivity, plus handy NLP utilities. Ideal for quick baselines and teaching.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

259

27 of 56

PRACTICAL

SENTIMENT APIS

Scoring a Batch with TextBlob

batch_sentiment.py

from textblob import TextBlob

import pandas as pd

df = pd.read_csv("tweets.csv")

def polarity(t):

return TextBlob(str(t)).sentiment.polarity

df["sentiment"] = df["text"].apply(polarity)

print(df.groupby(df["sentiment"] > 0)

.size())

Vectorise a column

apply() scores every row of a DataFrame in one line.

Aggregate

Group by polarity to summarise overall opinion.

Swap the backend

Replace TextBlob with a cloud API or transformer easily.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

260

28 of 56

DEEPER DIVE

SENTIMENT APIS

Local Models vs. Cloud APIs

Local / self-hosted

  • Full control & privacy
  • No per-call cost
  • Customisable to your domain
  • Needs setup and maintenance

Cloud API

  • Instant, production-grade
  • Multilingual and maintained
  • Pay per call; data leaves your box
  • Less domain control

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

261

29 of 56

DEEPER DIVE

SUPPORTING TOOLS

Data Handling: pandas & NumPy

pandas

DataFrames for loading, cleaning, joining, grouping and time-series analysis of social data. The backbone of every project in this course.

NumPy

Fast numerical arrays underpinning pandas, scikit-learn and matrix operations like similarity and factorisation.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

262

30 of 56

DEEPER DIVE

SUPPORTING TOOLS

Visualisation Libraries

matplotlib

The foundational plotting library — total control, publication-quality.

seaborn

Statistical charts with beautiful defaults on top of matplotlib.

Plotly

Interactive charts and dashboards for the web and notebooks.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

263

31 of 56

DEEPER DIVE

SUPPORTING TOOLS

Rounding Out the Stack

Gensim

Word2Vec, LDA and topic modelling at scale.

scikit-learn

Classical ML models, pipelines and metrics.

spaCy

Industrial-strength NLP: NER, POS, pipelines.

Hugging Face

Pre-trained transformers for text tasks.

Brandwatch / Hootsuite

Commercial social-listening suites.

Meltwater / Sprout

Enterprise monitoring and analytics.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

264

32 of 56

DEEPER DIVE

PUTTING IT TOGETHER

A Reference Analytics Architecture

1

Ingest

Tweepy / Scrapy collect data into raw storage.

2

Process

pandas + NLTK/spaCy clean and feature-engineer.

3

Analyse

NetworkX (structure) + scikit-learn (ML) generate insight.

4

Visualise

Gephi, matplotlib and Plotly present findings.

5

Report

Dashboards and notebooks communicate results.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

265

33 of 56

CORE SYLLABUS

HANDS-ON PROJECTS

Learning by Doing

THE PROJECT MINDSET

Analytics is a practical craft. Each hands-on project takes a real dataset through the full pipeline — collect, clean, analyse, visualise, interpret — building the skills the capstone will demand.

Real data

Work with messy, real-world social data.

Full pipeline

Practise every stage end to end.

Interpret

Turn output into a clear finding.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

266

34 of 56

PRACTICAL

HANDS-ON PROJECTS

Finding Real-World Datasets

Kaggle

Tweets, reviews, and labelled sentiment/fake-news sets.

SNAP

Large real network graphs for SNA.

Live APIs

Collect your own fresh data with Tweepy.

Open data portals

data.gov.in and civic datasets for context.

Paper datasets

Zenodo, figshare, Dataverse from publications.

Ethical scraping

Where permitted, build a bespoke dataset.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

267

35 of 56

PRACTICAL

MINI-PROJECT 1

Hashtag Trend Analysis

1

Collect

Tweepy pulls tweets for a hashtag over time.

2

Clean

Preprocess text; parse timestamps.

3

Analyse

Volume over time; top terms; sentiment.

4

Visualise

Time-series and word-frequency charts.

5

Interpret

When and why did it peak?

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

268

36 of 56

PRACTICAL

MINI-PROJECT 2

Community Mapping

1

Build graph

Retweet/mention edges in NetworkX.

2

Metrics

Centrality to find key accounts.

3

Detect

Louvain communities.

4

Visualise

Colour by community in Gephi.

5

Interpret

Who bridges the clusters?

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

269

37 of 56

PRACTICAL

MINI-PROJECT 3

Sentiment Dashboard

1

Collect

Brand mentions via API.

2

Score

VADER / model sentiment per post.

3

Aggregate

Trends by day, topic and region.

4

Dashboard

Interactive Plotly views.

5

Act

Surface issues and wins.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

270

38 of 56

DEEPER DIVE

GOOD PRACTICE

Reproducibility & Project Structure

Organise the repo

Separate data/, notebooks/, src/ and outputs/ with a clear README.

Pin dependencies

requirements.txt or an environment file so results reproduce.

Keep raw data immutable

Never overwrite raw data; derive cleaned copies.

Document decisions

Record queries, dates and cleaning choices for transparency.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

271

39 of 56

CORE SYLLABUS

CAPSTONE PROJECT

Analysing & Visualising Social-Media Trends

THE CAPSTONE

The capstone brings the whole course together: collect real social-media data, analyse it with network and machine-learning techniques, and visualise the trends you uncover — delivered as a working Python project and report.

Collect

Real data via API or dataset.

Analyse

SNA + ML from Modules III–IV.

Visualise

Clear, interpretable trend visuals.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

272

40 of 56

PRACTICAL

CAPSTONE PROJECT

Choosing a Topic

Trend tracking

How a hashtag or topic rises and falls over time.

Public opinion

Sentiment toward an event, brand or policy.

Community analysis

Structure of a conversation or fandom.

Influencer study

Who drives a topic and how far it spreads.

Misinformation

How a rumour propagates and who amplifies it.

Brand analytics

Competitive share of voice and perception.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

273

41 of 56

PRACTICAL

CAPSTONE PROJECT

Data-Collection Plan

1

Scope

Define topic, keywords, window.

2

Source

API, scrape or dataset.

3

Collect

Gather with provenance logged.

4

Validate

Check volume and quality.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

274

42 of 56

PRACTICAL

CAPSTONE PROJECT

Analysis Plan (SNA + ML)

Network analysis

Build the graph; compute centrality; detect communities; identify influencers and bridges.

Sentiment & text

Clean text; classify sentiment; extract topics and trends over time.

Diffusion

Trace how content spread; measure reach and cascade shape.

Quality checks

Filter bots and duplicates so findings reflect real activity.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

275

43 of 56

PRACTICAL

CAPSTONE PROJECT

Visualisation & Storytelling

Show the trend

Time-series of volume and sentiment make the story immediate.

Show the structure

A Gephi community map reveals who talks to whom.

Show the ranking

Bar charts of top influencers, terms and communities.

Tell the story

Lead with the insight; let each visual answer one question.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

276

44 of 56

PRACTICAL

CAPSTONE PROJECT

End-to-End Pipeline

1

Collect

Tweepy / dataset → raw CSV.

2

Preprocess

pandas + NLTK clean and structure.

3

Analyse

NetworkX + scikit-learn for structure and sentiment.

4

Visualise

matplotlib / Plotly / Gephi.

5

Report

Notebook + slides communicating findings.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

277

45 of 56

PRACTICAL

CAPSTONE PROJECT

A Capstone Skeleton

capstone.py

import pandas as pd, networkx as nx

from textblob import TextBlob

df = pd.read_csv("tweets.csv")

df["sent"] = df.text.apply(

lambda t: TextBlob(str(t)).sentiment.polarity)

# build mention graph

G = nx.from_pandas_edgelist(edges, "src", "dst")

pr = nx.pagerank(G)

daily = df.groupby(df.date)["sent"].mean()

daily.plot() # sentiment trend

Three stages, one script

Sentiment, network and trend in a compact skeleton.

Extend it

Add community detection, ML sentiment and bot filtering.

Then visualise

Export the graph to Gephi for the final figure.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

278

46 of 56

WORKED EXAMPLE

CAPSTONE PROJECT

Sample Result — A Sentiment Trend

The story at a glance

A Wednesday–Thursday dip flags a problem worth investigating.

Drill in

Filter to the dip and read the driving posts and terms.

Conclude

Tie the movement to a real event and recommend action.

Illustrative output — your capstone turns raw collection into a clear, defensible trend like this.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

279

47 of 56

CORE SYLLABUS

CAPSTONE PROJECT

Rubric & Deliverables

Component

What we assess

Weight

Data & collection

Sound, documented, ethical collection

20%

Analysis

Correct SNA + ML techniques applied

30%

Visualisation

Clear, honest, insightful visuals

20%

Interpretation

Meaningful, defensible conclusions

20%

Report & code

Reproducible, well-documented

10%

Deliverables: a Python project/notebook, visualisations, and a short report or presentation of findings.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

280

48 of 56

PRACTICAL

CAPSTONE PROJECT

A Suggested Timeline

Week

Focus

Milestone

1

Scope & question

Topic and plan approved

2

Data collection

Dataset gathered & validated

3

Preprocessing

Clean dataset ready

4

Network + ML analysis

Core results computed

5

Visualisation

Figures and dashboard

6

Report & present

Final submission

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

281

49 of 56

DEEPER DIVE

CAPSTONE PROJECT

Common Mistakes — Do & Don’t

Do

  • Pick a focused, answerable question
  • Document collection and cleaning
  • Validate findings against the raw data
  • Let visuals answer specific questions

Don’t

  • Collect data with no clear question
  • Over-clean and lose the signal
  • Report vanity metrics without context
  • Bury the insight under raw dumps

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

282

50 of 56

DEEPER DIVE

CAPSTONE PROJECT

Ethics & Compliance in Your Project

Respect the law

Follow the DPDP Act 2023, IT Act and platform terms in all collection and storage.

Anonymise

Remove personal identifiers; report aggregates, not individuals.

Secure your data

Store safely, limit access, and delete when the project ends.

Be transparent

State your data, methods and limitations honestly in the report.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

283

51 of 56

PRACTICAL

CAPSTONE PROJECT

Presenting & Documenting Your Work

Lead with the finding

Open with what you discovered, then show how.

One idea per visual

Each figure should answer a single, clear question.

State the method

Briefly note data, tools and techniques for credibility.

Acknowledge limits

Name sampling bias, bots and caveats — it builds trust.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

284

52 of 56

DEEPER DIVE

BEYOND THE COURSE

From Project to Portfolio & Career

Data analyst

Turn social data into business insight.

ML / NLP engineer

Build sentiment, recommender and detection models.

Network scientist

Study large-scale social and information networks.

Social-media analyst

Drive marketing and listening strategy.

Trust & safety

Fight misinformation, bots and abuse.

Researcher

Publish on computational social science.

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

285

53 of 56

A GUIDING PRINCIPLE

Tools serve questions, not the other way around. Choose the method because it answers your question — never run an algorithm just because you can.

NetworkX, Gephi, Tweepy, Scrapy and sentiment APIs are means to an end: understanding people and information through data.

286

54 of 56

DEEPER DIVE

QUICK REFERENCE

The Tool Landscape at a Glance

Stage

Tool

Purpose

Collect (API)

Tweepy

Pull data from X/Twitter

Collect (web)

Scrapy

Scrape permitted web data

Wrangle

pandas / NumPy

Clean and structure data

Analyse (graph)

NetworkX

Metrics, communities, links

Analyse (text)

scikit-learn / spaCy

ML and NLP

Sentiment

VADER / cloud APIs

Opinion scoring

Visualise

Gephi / Plotly

Networks and dashboards

Module V · Tools & Practical Implementation

CSIT406 · Social Media Analytics

287

55 of 56

MODULE V SUMMARY

Tools & Practice — Key Takeaways

NetworkX computes

The analytical engine for graphs and metrics.

Gephi reveals

Interactive visualisation of network structure.

Tweepy & Scrapy collect

API and web data, gathered responsibly.

Sentiment APIs score

Local or cloud opinion mining on demand.

Projects build skill

Real datasets, full pipeline, real insight.

The capstone ties it together

Collect → analyse → visualise trends in Python.

288

56 of 56

THANK YOU

Where graph theory meets

business intelligence.

From the structure of networks to the discipline of analytics — you now have the full toolkit to mine social media responsibly and well.

Dr. Indraneel Mukhopadhyay · Amity University Kolkata · CSIT406 Social Media Analytics

300