DEEP LEARNING FOR NATURAL LANGUAGE PROCESSING
- Academic year
- 2026/2027 Syllabus of previous years
- Official course title
- DEEP LEARNING FOR NATURAL LANGUAGE PROCESSING
- Course code
- CM0624 (AF:577098 AR:323992)
- Teaching language
- English
- Modality
- On campus classes
- ECTS credits
- 6
- Degree level
- Master's Degree Programme (DM270)
- Academic Discipline
- INF/01
- Period
- 1st Semester
- Course year
- 2
- Where
- VENEZIA
- Moodle
- Go to Moodle page
Contribution of the course to the overall degree programme goals
The programme is organized into five blocks.
Block 1 — Foundations of NLP and Statistical Modeling
The first block introduces the main problems of Natural Language Processing and the mathematical and computational tools required to address them.
Topics include:
- introduction to NLP and Large Language Models;
- mathematical and machine-learning foundations;
- text processing and tokenization;
- logistic regression and text classification;
- n-gram language models.
Block 2 — Neural Representations and Sequence Models
The second block introduces learned representations and neural models for processing linguistic sequences.
Topics include:
- feed-forward neural language models;
- distributed word representations;
- NLP programming with PyTorch;
- recurrent neural networks;
- sequence-to-sequence models;
- classical attention and self-attention.
Block 3 — Transformers and Generative Modeling
The third block presents the Transformer architecture and the main families of pretrained language models.
Topics include:
- Transformer architecture;
- encoder-only, decoder-only, and encoder–decoder configurations;
- pretraining objectives;
- language generation;
- text-to-text models, including T5;
- fine-tuning Transformers for structured and generative NLP tasks.
Block 4 — Adapting and Extending Language Models
The fourth block examines the methods used to transform a pretrained model into a complete NLP system.
Topics include:
- sentence representations and semantic retrieval;
- Retrieval-Augmented Generation;
- post-training and alignment;
- scaling and long-context models;
- analysis of Transformer parameters;
- Parameter-Efficient Fine-Tuning;
- implementation of a Llama-inspired decoder-only Transformer.
Block 5 — Reasoning, Agents, Evaluation, and Security
The final block examines the capabilities, applications, and risks of modern LLM-based systems.
Topics include:
- reasoning models;
- LLM agents and external tools;
- evaluation of NLP systems and LLMs;
- language-model security;
- privacy and memorization of training data.
Expected learning outcomes
- understand the linguistic, mathematical and computational foundations of NLP;
- describe the development from statistical language models to neural models and Transformers;
- process and tokenize text collections;
- implement and apply algorithms for text classification, representation and generation;
- understand and apply neural networks, recurrent models, sequence-to-sequence models and attention mechanisms;
- understand the Transformer architecture and the main families of language models;
- use PyTorch and Transformers to develop and train NLP models;
- apply full fine-tuning and parameter-efficient fine-tuning techniques;
- build systems for sentence representation, semantic retrieval and RAG;
- select appropriate metrics and protocols for evaluating NLP systems and LLMs;
- understand the principles of post-training, alignment, scaling and long-context modelling;
- critically assess data quality, provenance and representativeness, as well as the risks, limitations and social impact of NLP systems;
- select models and methodologies appropriate to the requirements of a specific application.
Pre-requirements
Basic Python programming skills are required for the practical activities. Previous familiarity with PyTorch and the Transformers library is useful but not mandatory.
Contents
2 Primer Mathematical and Learning Foundations + Neural Networks intro
3 Lecture Text Processing and Tokenization
4 Lecture Logistic Regression and Text Classification
5 Lecture N-gram Language Models
6 Lecture Neural Language Models (and RNN LM ?)
7 Lecture Distributed Word Representations
8 Lab PyTorch NLP Lab
9 Lecture RNNs, Seq2Seq and Classical Attention
10 Lecture Self-Attention
11 Lecture Transformer Architecture
12 Lecture Pretraining Objectives and Model Families
13 Lecture Language Generation + T5
14 Lab Transformers Coding Lab - Structured and Text-to-Text NLP (fine-tuning)
15 Lecture Sentence Representations and Retrieval
16 Lecture Retrieval-Augmented Generation
17 Lecture Post-Training and Alignment
18 Lecture Scaling and Long-Context Models
19 Lecture Transformer Parameters and Parameter Efficient Fine-Tuning
20 Lab Building a Decoder-Only Transformer Model Like Llama
21 Lecture Reasoning models
22 Lecture LLM agents and Tools
23 Lecture Evaluation of NLP Systems and LLMs
24 Lecture LLM security & memorization
Referral texts
Daniel Jurafsky and James H. Martin, Speech and Language Processing, 3rd edition draft. https://web.stanford.edu/~jurafsky/slp3/
Lecture slides, scientific papers, notebooks, code examples, and other required materials will be made available through Moodle. Additional resources may be recommended during the course.
Assessment methods
Students may optionally undertake an individual Python project agreed upon in advance with the instructor. The project will be assessed according to its methodological and technical quality, analysis of results and presentation, and may contribute up to 3 points to the final grade. The project does not replace the oral examination, and points will be awarded only if the examination is passed.
The examination is oral and covers the entire course syllabus. Students must demonstrate knowledge of the theoretical concepts, models, algorithms and methodologies addressed during the course.
Any optional individual project will be presented and discussed during the examination and may contribute up to 3 additional points, subject to the maximum grade permitted.
The instructor is responsible for ensuring the authenticity and originality of the assessment. In the event of suspected irregularities, an additional assessment may be required, including through a different format.
Type of exam
The instructor is responsible for ensuring the authenticity and originality of all examinations and coursework. In cases of suspected academic misconduct, an additional on-site assessment may be required during the exams, which may differ from the standard format.
Grading scale
18–22: sufficient knowledge, basic application skills and adequate technical language.
23–26: fair knowledge, good ability to connect and apply concepts, and adequate critical thinking.
27–30: in-depth knowledge, autonomous application of concepts, and excellent critical and presentation skills.
Honours: excellent knowledge, full autonomy of judgement and the ability to produce original insights.
Teaching methods
- lectures covering theoretical and methodological foundations;
- discussion of examples, case studies and research papers;
- guided exercises;
- Python programming laboratories;
- activities involving the design, implementation and evaluation of NLP systems.
The practical activities will allow students to consolidate the theoretical concepts through the use of PyTorch and Transformers and through the implementation of fundamental components of modern language models.