Skip to content

Sentiment Analysis of Telugu–English Code-Mixed Social-Media Text: Multilingual Transformers vs BiLSTM Baselines

  • 13 slides
  • 15 viva questions
  • 5 modules
  • Code included

@telugu-english-code-mixed-sentimentUpdated Oct 2026

A new annotated Telugu–English corpus with measured agreement, and an honest benchmark of TF-IDF+SVM, BiLSTM, mBERT, XLM-R and MuRIL

M.Tech / M.E., Artificial Intelligence & Machine Learning · Sem 4 · Advanced · 24 weeks · Solo

More info
Level
Advanced · 24 weeks · Solo
Relevant for
Telangana
Common at
JNTU Hyderabad, Anna University, Visvesvaraya Technological University
Syllabus
JNTUH M.Tech R22 · Dissertation Work Review-III + Dissertation Viva-Voce · Semester 4
Tech stack
  • Python
  • PyTorch
  • Hugging Face Transformers (industry-standard)
  • scikit-learn
  • FastText
  • pandas
  • Label Studio (annotation)
For educational purposes only

Unlock this project

Full PPT + speaker notes, source code and setup steps, READMEFIRST, instructions and all 15 viva answers.

One-time. No subscription, no auto-renew, no drama.

Project packs

Credits never expire and work on any project. Use one here, save the rest for your friend who “will pay you back”.

  1. Pinned

    1 min

    Overview

    Telugu speakers on social media rarely write in one language or one script. A typical comment under a Telugu film trailer or a Hyderabad news video mixes Romanised Telugu, English words and occasionally Telugu script — "movie asalu bagaledu, waste of time" or "BGM keka 🔥". Standard sentiment tools, trained on monolingual English or on Telugu-script text, handle such input poorly, and publicly available labelled Telugu–English code-mixed data is scarce compared with Tamil, Malayalam or Hindi.

    This dissertation does two things. First, it constructs a sentiment corpus of public Telugu–English code-mixed comments with written annotation guidelines, two independent annotators and inter-annotator agreement measured by Cohen's kappa. Second, it benchmarks classical and neural approaches on that corpus: TF-IDF with a linear SVM, a BiLSTM with FastText subword embeddings, and fine-tuned pretrained multilingual transformers — multilingual BERT, XLM-RoBERTa and MuRIL — using macro-F1 over repeated runs, followed by a structured error analysis by code-mixing level, script, negation and sarcasm.

    The work applies the Deep Learning and NLP electives of the JNTUH R22 M.Tech programme and is implemented in Python with PyTorch and scikit-learn; Hugging Face Transformers is used as an industry-standard extra. Ethics are built in: only public comments are collected through official interfaces, user handles are removed, and no personal data is stored. Results are produced by the student's own experiments — this bundle provides the protocol and empty results tables.

    Syllabus alignment

    JNTUH · M.Tech R22

    Dissertation Work Review-III + Dissertation Viva-Voce · Semester 4 · 2 credits

    Subjects this project applies
    • Natural Language Processing (elective)
    • Deep Learning (elective)
    • Advanced Data Structures Lab
    • Dissertation Work Review-II (Sem 3: literature and corpus design)
    How it is evaluated

    See your department's project guidelines.

    Also fits: Anna University M.E. Regulation 2021, VTU M.Tech 2022 Scheme.

    1 min read · 15 viva questions

  2. 2 min

    Synopsis

    Abstract

    Code-mixed text — language alternation within a sentence — is the norm in Indian social media, but sentiment resources for Telugu–English remain limited. This work builds an annotated corpus of Telugu–English code-mixed comments (target 12,000 items, three sentiment classes plus a not-Telugu filter label) with documented guidelines and measured inter-annotator agreement, and compares five models: TF-IDF + linear SVM, BiLSTM with FastText embeddings, multilingual BERT, XLM-RoBERTa base and MuRIL. Models are evaluated with macro-F1 over five seeds on a fixed test split, with statistical significance testing and error analysis stratified by code-mixing index and script.

    Introduction

    Pretrained multilingual encoders have changed multilingual NLP, but most are pretrained on native-script text, while Telugu users frequently type in Roman script with inconsistent spellings ("bagundi", "baagundhi", "bgndi"). Whether such encoders transfer to Romanised, code-mixed Telugu — and whether an Indian-language-focused encoder like MuRIL has an advantage — is an empirical question this dissertation answers on a carefully constructed corpus.

    Literature Gap

    • Shared tasks and datasets for Dravidian code-mixed sentiment (e.g. DravidianCodeMix) cover Tamil, Malayalam and Kannada; Telugu–English resources are fewer and smaller.
    • Many studies report accuracy without agreement statistics for their labels, which makes it hard to know the ceiling.
    • Gap: a documented Telugu–English sentiment corpus with measured agreement, and a controlled benchmark of classical, recurrent and transformer models with transliteration-handling ablations.

    Proposed Work

    • Corpus construction pipeline: collection via official APIs, filtering, language/script tagging, anonymisation, annotation in Label Studio, adjudication.
    • Transliteration strategies: raw text, Roman-normalised, and Telugu-script back-transliteration, compared as an ablation.
    • Benchmark and error analysis with significance testing.

    Feasibility

    All tools are open-source; fine-tuning base-size encoders on about 10,000 short comments needs a single GPU with 12–16 GB memory (college lab GPU or a cloud notebook). Annotation of 12,000 items by two annotators is achievable in about five weeks at a few hundred items per day each.

  3. 1 min

    Problem statement

    Sentiment analysis tools used by businesses, film producers, political researchers and public-service departments in Telangana and Andhra Pradesh fail on the way people actually write online: Telugu and English mixed within a sentence, mostly in Roman script with non-standard spelling, plus emojis and slang. Monolingual models misread such text, and the lack of a documented, agreement-checked Telugu–English sentiment corpus makes it impossible to measure how well modern pretrained multilingual transformers actually perform compared with cheaper classical and recurrent baselines.

    This dissertation addresses two questions: (1) Can a reliable Telugu–English code-mixed sentiment corpus be built from public social-media comments with substantial inter-annotator agreement? and (2) On that corpus, how do TF-IDF + SVM, BiLSTM with FastText, mBERT, XLM-R and MuRIL compare in macro-F1, and which linguistic phenomena — code-mixing level, script, negation, sarcasm — account for their errors?

  4. 1 min

    Objectives & scope

    1. 01Collect public Telugu–English code-mixed comments through official platform interfaces and filter them by language and script.
    2. 02Write annotation guidelines and label a corpus of about 12,000 comments with two annotators, reporting Cohen's kappa before and after guideline revision.
    3. 03Anonymise the corpus by removing user handles, URLs, phone numbers and e-mail addresses, and prepare a data card.
    4. 04Implement TF-IDF + linear SVM and BiLSTM + FastText baselines.
    5. 05Fine-tune mBERT, XLM-RoBERTa base and MuRIL with a common protocol and compare macro-F1 over five seeds with significance tests.
    6. 06Run an ablation on transliteration handling and a structured error analysis by code-mixing index, script, negation and sarcasm.
    7. 07Document the work as a dissertation and a conference-style paper.

    Scope

    In scope

    • Public comments on Telugu film, sports and news content, collected through official APIs within their terms of service.
    • Three sentiment classes (positive, negative, neutral) plus a not-Telugu label used only for filtering.
    • Classical, recurrent and base-size transformer models; transliteration ablation; error analysis.
    • Ethics: public data only, anonymisation, no attempt to identify or profile individuals.

    Out of scope

    • Hate-speech or offensive-language classification (separate task and ethics review).
    • Large encoder variants and generative models; speech or image content.
    • Releasing raw comment text where platform terms forbid it (only IDs and labels would be shared).
  5. 2 min

    Methodology

    Research design: a two-part empirical study — (A) corpus construction with a reliability study, and (B) a controlled comparative experiment with repeated runs.

    StageWeeksActivitiesOutput
    Review-II recap1–2Finalise research questions, guidelines draft, pilot 300 itemsPilot kappa, revised guidelines
    Collection & cleaning3–5API collection, deduplication, language/script filtering, anonymisationRaw corpus + data card
    Annotation6–10Two annotators, weekly calibration, adjudication of disagreementsLabelled corpus, kappa report
    Baselines10–12TF-IDF + SVM, BiLSTM + FastTextBaseline table
    Transformers13–16mBERT, XLM-R, MuRIL fine-tuning, 5 seeds eachMain results table
    Ablations & error analysis17–19Transliteration variants, CMI buckets, phenomenon taggingAnalysis chapter
    Writing & Review-III20–24Dissertation, paper, plagiarism check, viva preparationSubmission

    Annotation protocol

    • Guidelines define each class with 10+ Telugu–English examples, rules for mixed sentiment (label the dominant sentiment; if balanced, neutral), sarcasm (label intended sentiment) and emojis.
    • Each item labelled independently by two annotators; Cohen's kappa reported per batch. Target: κ ≥ 0.61 ("substantial" in the Landis–Koch scale); if a batch falls below, guidelines are revised and the batch re-labelled.
    • Disagreements are adjudicated by a third person (the guide or a senior student) and logged.

    Experimental protocol

    • Stratified 70/15/15 split, fixed once; no near-duplicate comments across splits (MinHash deduplication).
    • Transformers: learning rate 2e-5 to 5e-5, batch 16–32, up to 5 epochs, early stopping on validation macro-F1, max length 128.
    • Metrics: macro-F1 (primary), per-class F1, accuracy; mean ± sd over 5 seeds; paired bootstrap or McNemar's test between the best transformer and best baseline.
    • Hypotheses (not results): transformers exceed the best baseline by ≥ 3 macro-F1 points; MuRIL is at least as good as XLM-R on Romanised Telugu; errors concentrate in high code-mixing and sarcastic items.
  6. 2 min

    Architecture & tech stack

    • Python
    • PyTorch
    • Hugging Face Transformers (industry-standard)
    • scikit-learn
    • FastText
    • pandas
    • Label Studio (annotation)

    The pipeline has three stages: corpus construction, modelling, and evaluation and analysis.

    flowchart TD
      A["Public comments via official APIs"] --> B["Deduplicate and filter by language and script"]
      B --> C["Anonymise: remove handles, URLs, phone numbers"]
      C --> D["Annotation in Label Studio: 2 annotators"]
      D --> E{"Cohen's kappa at least 0.61?"}
      E -- No --> F["Revise guidelines and re-label batch"]
      F --> D
      E -- Yes --> G["Adjudicate disagreements"]
      G --> H["Stratified train, validation, test split"]
      H --> I["TF-IDF + linear SVM"]
      H --> J["BiLSTM + FastText embeddings"]
      H --> K["mBERT, XLM-R, MuRIL fine-tuning"]
      I --> L["Evaluator: macro-F1 over 5 seeds"]
      J --> L
      K --> L
      L --> M["Significance tests and error analysis"]

    Text preprocessing variants (ablation)

    • Raw: only lower-casing of Roman text, emoji kept as tokens.
    • Roman-normalised: rule-based spelling normalisation for frequent Telugu words (e.g. collapsing repeated vowels, common suffix variants).
    • Script-unified: Romanised Telugu tokens back-transliterated to Telugu script with a public transliteration tool, English tokens left as-is.

    Models

    • TF-IDF + LinearSVC: word 1–2-grams plus character 2–5-grams (robust to spelling variation); C tuned on validation.
    • BiLSTM: FastText skip-gram embeddings (300-d, subword n-grams) trained on the unlabelled part of the corpus; 2-layer BiLSTM (128 units), max-pool, dropout 0.3.
    • Transformers: bert-base-multilingual-cased, xlm-roberta-base, google/muril-base-cased with a linear classification head on the [CLS]/<s> representation.

    Results tables (to be filled from your runs)

    ModelPreprocessingMacro-F1 (mean ± sd)F1 posF1 negF1 neu
    TF-IDF + SVMraw
    BiLSTM + FastTextraw
    mBERTraw
    XLM-R baseraw
    MuRILraw
    Best modelscript-unified
    Annotation batchItemsCohen's κGuideline version
    Pilot300v1
    Batch 13,000
  7. 5 modules

    Modules

    • Corpus Collection and Anonymisation

      Collects public comments through official APIs with rate limiting, stores only comment ID, text and timestamp, removes duplicates and non-Telugu items using a script and keyword heuristic, and strips user handles, URLs, phone numbers and e-mail addresses with tested regular expressions.

    • Annotation and Agreement

      Label Studio project configuration, versioned annotation guidelines, batch export scripts, Cohen's kappa computation per batch with confusion tables between annotators, and an adjudication log recording the final label and reason for every disagreement.

    • Classical and Recurrent Baselines

      Scikit-learn pipeline for TF-IDF word and character n-grams with LinearSVC and grid search, plus a PyTorch BiLSTM classifier using FastText embeddings trained on the unlabelled corpus, with early stopping and class-weighted loss.

    • Transformer Fine-tuning

      A common training script for mBERT, XLM-R and MuRIL using Hugging Face Transformers with PyTorch, configurable learning rate, sequence length and seeds, mixed-precision training, and checkpointing of the best validation macro-F1 model.

    • Evaluation and Error Analysis

      Aggregates metrics across seeds, runs paired bootstrap and McNemar tests, computes a code-mixing index per comment, buckets errors by CMI, script, negation and sarcasm tags, and produces confusion matrices and example tables for the dissertation.

  8. Locked

    Presentation

    13 slides with speaker notes. The outline below is free; the bullets, notes and the generated .pptx unlock with the project.

    1. Telugu–English Code-Mixed Sentiment Analysis
    2. Motivation
    3. Research Questions
    4. Literature Review
    5. Corpus Construction
    6. Annotation Guidelines and Agreement
    7. Corpus Statistics
    8. Models
    9. Experimental Protocol
    10. Results
    11. Error Analysis
    12. Ethics and Limitations
    13. Conclusion and Future Work

    Bullets, speaker notes and the .pptx download unlock with the project.

    Presentation is locked: 13 slides, Speaker notes, .pptx download.

  9. 1 min

    Future scope

    • Aspect-based sentiment for film reviews (story, music, acting) in code-mixed text.
    • Distillation and quantisation of the best model for mobile or CPU deployment.
    • Cross-lingual transfer from Telugu to Kannada–English and Tamil–English corpora.
    • Semi-supervised learning using the large unlabelled pool with pseudo-labels.
    • Offensive-language detection as a separate, ethics-reviewed extension.
  10. 10 sources

    References

    1. J. Devlin, M.-W. Chang, K. Lee & K. Toutanova, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, NAACL 2019
    2. A. Conneau et al., Unsupervised Cross-lingual Representation Learning at Scale (XLM-R), ACL 2020
    3. S. Khanuja et al., MuRIL: Multilingual Representations for Indian Languages, 2021
    4. P. Bojanowski, E. Grave, A. Joulin & T. Mikolov, Enriching Word Vectors with Subword Information, TACL 2017
    5. J. Cohen, A Coefficient of Agreement for Nominal Scales, Educational and Psychological Measurement, 1960
    6. J. R. Landis & G. G. Koch, The Measurement of Observer Agreement for Categorical Data, Biometrics, 1977
    7. B. R. Chakravarthi et al., DravidianCodeMix: Sentiment Analysis and Offensive Language Identification Dataset for Dravidian Languages in Code-Mixed Text, Language Resources and Evaluation, 2022
    8. Hugging Face Transformers Documentation
    9. PyTorch Documentation
    10. D. Jurafsky & J. H. Martin, Speech and Language Processing, 3rd ed. draft

    Cite this bundle

    OnlyProjects. (2026). Sentiment Analysis of Telugu–English Code-Mixed Social-Media Text: Multilingual Transformers vs BiLSTM Baselines: M.Tech / M.E. Artificial Intelligence & Machine Learning project bundle [Educational resource]. https://onlyprojects.online/projects/mtech-aiml-telugu-english-code-mixed-sentiment

Slides, diagrams & files

13 slides. Titles are free; bullets, speaker notes and the .pptx unlock with the project.

  1. SLIDE 1

    Telugu–English Code-Mixed Sentiment Analysis

  2. SLIDE 2

    Motivation

  3. SLIDE 3

    Research Questions

  4. SLIDE 4

    Literature Review

  5. SLIDE 5

    Corpus Construction

  6. SLIDE 6

    Annotation Guidelines and Agreement

  7. SLIDE 7

    Corpus Statistics

  8. SLIDE 8

    Models

  9. SLIDE 9

    Experimental Protocol

  10. SLIDE 10

    Results

  11. SLIDE 11

    Error Analysis

  12. SLIDE 12

    Ethics and Limitations

  13. SLIDE 13

    Conclusion and Future Work

Architecture diagram

1
flowchart TD
  A["Public comments via official APIs"] --> B["Deduplicate and filter by language and script"]
  B --> C["Anonymise: remove handles, URLs, phone numbers"]
  C --> D["Annotation in Label Studio: 2 annotators"]
  D --> E{"Cohen's kappa at least 0.61?"}
  E -- No --> F["Revise guidelines and re-label batch"]
  F --> D
  E -- Yes --> G["Adjudicate disagreements"]
  G --> H["Stratified train, validation, test split"]
  H --> I["TF-IDF + linear SVM"]
  H --> J["BiLSTM + FastText embeddings"]
  H --> K["mBERT, XLM-R, MuRIL fine-tuning"]
  I --> L["Evaluator: macro-F1 over 5 seeds"]
  J --> L
  K --> L
  L --> M["Significance tests and error analysis"]

Files

Viva questions & answers

3 of 15 questions free. Explain each answer in your own words before you move on.

  1. General

    What is the research contribution of your dissertation?

    Two contributions: a documented Telugu–English code-mixed sentiment corpus with guidelines and measured Cohen's kappa, and a controlled benchmark of TF-IDF + SVM, BiLSTM + FastText, mBERT, XLM-R and MuRIL with transliteration ablations and error analysis stratified by code-mixing level.

  2. General

    What did you present at Review-II and what is new at Review-III?

    Review-II covered the literature review, research questions, guideline draft, pilot annotation with its first kappa and the collection pipeline. Review-III adds the full labelled corpus, all baseline and transformer experiments, ablations, error analysis, the dissertation and a paper draft.

  3. Concept

    What is code-mixing and how is it different from code-switching?

    Code-mixing usually means alternating languages within a sentence or phrase, such as Telugu verbs with English nouns, while code-switching often refers to switching between sentences. Our comments show both, mostly intra-sentential mixing written in Roman script.

+12 more questions

They and the answers unlock with the project. Try answering the ones above yourself first. Your examiner will.

For educational purposes only. Use this bundle to understand how the project works, then build and write your own. Submitting it verbatim is between you, your conscience and your external examiner.