Welcome
Welcome to the Introduction to Deep Learning course website. Here you will find all the necessary information related to our class.
Class Schedule
Classes are held Tuesdays at 16:15. Dates and rooms are listed below.
- Class 1: 22/09/2026, 16:15–18:15, room S207. CM: Introduction, general principles, datasets, simple example
- Class 2: 29/09/2026, 16:15–18:45, room S206. TD 1
- Class 3: 06/10/2026, 16:15–18:15, room S308. CM: Gradient descent and backpropagation
- Class 4: 13/10/2026, 16:15–18:45, room S206. TD 2, followed by QCM 1
- Class 5: 27/10/2026, 16:15–18:15, room S208. CM: Convolutional neural networks and autoencoders
- Class 6: 03/11/2026, 16:15–18:45, room P122. TD 3, followed by QCM 2
- Class 7: 10/11/2026, 16:15–18:15, room P122. CM: Outlook and advanced topics
- Class 8: 17/11/2026, 16:15–18:45, room S206. TD 4
- Class 9: 24/11/2026, 16:15–18:15, room S306. Project work
- Class 10: 01/12/2026, 16:15–18:15, room S206. Presentations and project discussions
Mini-projects
Projects are to be carried out in groups of two or three. You can either select a project from the ideas below or suggest your own:Datasets (select one)
- The "Fashion MNIST" dataset is a collection of 60,000 training and 10,000 test grayscale images of 10 different fashion categories, each 28x28 pixels in size. Designed as a more challenging alternative to the classic MNIST dataset of handwritten digits, it includes images of items like t-shirts, trousers, pullovers, dresses, and shoes.
- The "Flickr-Faces-HQ (FFHQ)" dataset consists of 70,000 high-quality PNG images (for our purposes, use the 128x128 pixel thumbnails) of human faces, each at a resolution of 1024x1024 pixels, sourced from Flickr. Designed to provide a diverse set of real-world face images, the dataset spans a wide range of ages, ethnicities, and image backgrounds.
- The DeepSat (SAT-4) and (SAT-6) Airborne Image Classification datasets each comprise about half a million 28x28 pixel multispectral satellite images, representing four/six distinct land cover types. Sourced from airborne sensors, these datasets offer a rich collection of images to facilitate the development and benchmarking of classification models for aerial earth observation tasks.
- The CIFAR-10 dataset consists of 60,000 32x32 color images spanning 10 different classes, with 6,000 images per class. The classes include objects like airplanes, cars, birds, cats, deer, dogs, frogs, horses, ships, and trucks.
Tasks (select one, but beware: not all tasks are compatible with all data sets)
- Train a convolutional neural network model to accurately categorize images into predefined classes. The model will learn distinctive features and patterns for each category. Try to optimize for classification accuracy.
- Design and train a neural network that determines the orientation of images, categorizing them as upright, sideways, or upside-down. To create the dataset, start with upright images and programmatically rotate them to various orientations, ensuring each image has a corresponding labeled orientation.
- Design and train a neural network model that can accurately predict the correct order of a set of four subimages, which when arranged correctly, create a coherent larger image. The dataset can be created by taking complete images and segmenting them into a 2x2 grid of subimages. After shuffling the order of these subimages, the correct sequence will be recorded. To represent the 24 possible orderings, a one-hot encoding scheme may be employed.
- Train a denoising autoencoder, a specialized type of neural network, to remove noise from digital images. By training on pairs of noisy and clean images, the model learns to reconstruct high-quality, noise-free visuals. The outcome will be a powerful tool that enhances image clarity, surpassing traditional denoising techniques. You will need to artificially add noise in order to create the noisy of the database.
- Implement a variational autoencoder (VAE), a specialized type of deep learning model, to learn compressed, latent representations of images. By encoding high-dimensional input data into a lower-dimensional latent space and subsequently decoding it back, the VAE will be trained to both compress and reconstruct images with minimal loss of information. The outcome will be a model that enables image generation.
| Dataset/Task | Simple classification | Image orientation | Subimage permutation | Denoising auto-encoder | Super-resolution auto-encoder |
|---|---|---|---|---|---|
| MNIST | Dataset | Dataset | Dataset | Dataset | Dataset |
| Fashion MNIST | Dataset | Dataset | Dataset | Dataset | Dataset |
| Flickr-Faces-HQ (FFHQ) | NA | Dataset | Dataset | Dataset | Dataset |
| SAT-4 | Dataset | NA | NA | Dataset | Dataset |
| SAT-6 | Dataset | NA | NA | Dataset | Dataset |
| CIFAR-10 | Dataset | Dataset | Dataset | Dataset | Dataset |
| Letter dataset | Dataset | Dataset | Dataset | Dataset | Dataset |
- Fraud detection dataset: Dataset
Grades
- Two multiple-choice tests (75%)
- Participation in the project discussions (25%)
- Please share your feedback about the module anonymously: Feedback form
Textbooks
The course is inspired by the following books, which are all recommended:- Simon J.D. Prince "Understanding Deep Learning", MIT Press (2023)
- Ian Goodfellow and Yoshua Bengio and Aaron Courville: "Deep Learning", MIT Press (2016)
- François Chollet: "Deep Learning with Python", Second Edition, Manning (2021)
Important Papers
- Perceptron (Rosenblatt, 1957)
- Dense Neural Network (LeCun, 1989)
- Convolutional Neural Network (LeCun et al., 1998)
- Variational autoencoder (Kingma & Welling, 2013)
- Adam optimizer (Kingma & Ba, 2014)
- Weight initialization (He et al., 2015)
- Batch normalization (Ioffe & Szegedy, 2015)
- Residual connections (He et al., 2015)
- Layer normalization (Ba et al., 2016)
- Transformer architecture (Vaswani et al., 2017)
Contact
Daniel FORSTER
Email: daniel.forster[at]univ-orleans.fr