STATISTICAL MODELING

Academic year
2026/2027 Syllabus of previous years
Official course title
STATISTICAL MODELING
Course code
CT0676 (AF:522195 AR:301196)
Teaching language
English
Modality
On campus classes
ECTS credits
6
Degree level
Bachelor's Degree Programme
Academic Discipline
SECS-S/01
Period
1st Semester
Course year
3
Where
VENEZIA
Moodle
Go to Moodle page
This course belongs to the interdisciplinary educational activities of the Data Science curriculum of the Bachelor in Informatics. The course is designed to give a panoramic view of selected statistical modelling methods, at an intermediate level.
The course covers the main concepts in linear models and generalized linear models and possibly further extension of these modelling frameworks including time series analysis. The focus is placed on providing the main insights on the statistical/mathematical foundations of the models and on showing the effective implementation of the methods through the use of statistical software. This is achieved by a mixture of theory and reproducible code. Real data examples and case studies are also introduced.
1. Knowledge and understanding
- Know and understand the mathematical concepts underlying linear models, including generalized models, and their estimation
- Understand the relationship between linear models and basic probabilistic and inferential concepts
- Know and understand the different types of statistical analyses with predictive purpose for which linear models, including generalized models, can be used

2. Ability to apply knowledge and understanding

- Ability to apply techniques for analyzing, designing and solving statistical problems.
- Ability to apply data processing techniques to real data using appropriate statistical software.
- Ability to use the results of statistical model estimation in the context of prediction and classification

3. Communication Skills.

- Being able to use technical language and notation to communicate the details of a predictive statistical model.
- Being able to translate statistical and mathematical concepts related to linear models into common language and vice versa.
Students are assumed to have reached the learning objectives of the courses
Calculus 1 and 2
Linear Algebra
Probability and Statistics (formerly Probability and Statistics and Data Analysis) although it is not formally required to have passed the examination.
1. Introduction
1.1 Course overview
1.2 What is predictive modeling?
1.3 General notation and background

2. Linear models I: simple and multiple linear model
2.1 Model formulation and least squares
2.2 Assumptions of the model
2.3 Inference for model parameters
2.4 Prediction
2.5 ANOVA
2.6 Model fit

3. Linear models II: model selection, extensions, and diagnostics
3.1 Model selection
3.2 Use of qualitative predictors
3.3 Nonlinear relationships
3.4 Model diagnostics
3.5 Potential critical issues in regression models

4. Generalized linear models
4.1 Model formulation and estimation
4.2 Inference for model parameters
4.3 Prediction
4.4 Deviance
4.5 Model selection
4.6 Classification for binary data

If time allows:
5. Time-series forecasting
5.1: elements of time series
5.2: auto-regressive models
5.3: exponential smoothing forecasting

The program might be slightly modified during the semester. Students are encouraged to actively request for the course to also cover specific statistical questions of interest.
Julian J. Faraway, 2014. Linear Models with R Second Edition, Chapman and Hall/CRC
Julian J. Faraway, 2016. Extending the Linear Model with R: Generalized Linear, Mixed Effects and Nonparametric Regression Models, Second Edition Chapman and Hall/CRC
Peter H. Westfall, Andrea L. Arias, Understanding Regression Analysis - A Conditional Distribution Approach, Chapman and Hall/CRC
James, Gareth, Daniela Witten, Trevor Hastie, and Robert Tibshirani. 2023. An Introduction to Statistical Learning (second edition) . Springer (A free copy is here https://www.statlearning.com/ )
The exam lasts 90 minutes, takes place in the IT lab and is composed of two parts: a written part and an R-based part for a total of 33 points. Both parts are composed of exercises which aim to evaluate
1. the theoretical knowledge of the course topics,
2. the ability to apply them for solving real data problems,
(1+2) max 20 points
3. the ability to use R and interpret its output to solve real data problems,
4. the ability to use the R software to present the results of a statistical data analysis.
(3+4) max 13 points
The instructor is responsible for ensuring the authenticity and originality of all exams and assignments completed during the course. In the event of suspected academic misconduct, an additional in-person assessment may be required after the exams, which may differ from the standard format.
written

The instructor is responsible for ensuring the authenticity and originality of all examinations and coursework. In cases of suspected academic misconduct, an additional on-site assessment may be required during the exams, which may differ from the standard format.

Typically the grading will follow the following criteria:
A. grades between 18 and 22 will be assigned when there is evidence of
- comprehension of basic theoretical concepts underlying statistical modelling;
- limited ability to interpret and present a statistical modelling;
- limited ability to adapt the statistical modelling to the problem address in a specific real case;
B. grades between 23 and 26 will be assigned when there is evidence of
- comprehension of theoretical concepts underlying statistical modelling beyond the basic ones;
- moderate ability to interpret and present a statistical modelling
- moderate ability to adapt the statistical modelling to the problem address in a specific real case;
C. grades between 27 and 30 will be assigned when there is evidence of
- good or very good comprehension of theoretical concepts underlying statistical modelling;
- good or very good ability to interpret and present a statistical modelling;
- good or very good ability to adapt the statistical modelling to the problem address in a specific real case;
D. Honors (lode) will be awarded to students who demonstrate a particular ability to answer all questions thoroughly, with attention to detail and care in their written work.
This course is based on lectures, which will cover the major topics, emphasizing and discussing the important points. Theoretical lectures will be complemented by exercise classes and lab sessions. The statistical software used in the course is R (www.r-project.org). The personal participation is important, and it will help the student to learn more efficiently to read the assigned material to reinforce the lectures.
Access to Moodle for the 2026-2027 academic year will be granted to students who have the course in their study plan. If you would like to have access to the materials, please send a request to the teacher, specifying your name, surname, student ID number, degree program and attaching a screenshot of your study plan.
Definitive programme.
Last update of the programme: 19/09/2026