Understanding Deep Learning
Introduction
Deep Learning is a subset of machine learning, which is designed in various neural layers in a way to mimic human brain. It is a powerful machine learning technique, which uses multiple layers of neural network and works with large data and is trained to recognize patterns, pictures, speech and languages, with high accuracy.

Image Source: https://www.xavyi.site/?ggcid=2594508
AI models are trained to uncover underlying patterns, build models and make predictions. Deep Learning has gained academic and industrial interests from the early 2000s.
Machine Learning vs Deep Learning
As already mentioned, deep learning is a subfield of machine learning.
· Machine Learning applies statistical models to learn patterns from data, while Deep Learning uses artificial neural networks.
· ML is better with small or medium data and is used for simple tasks, whereas DL requires large data and is used for complex tasks like image, text and language processing.
· ML needs human intervention for training models, while DL eliminates most of the human intervention and is capable of self training with the use of neural networks.
· ML is less complex and easy to understand. On the other hand, DL is more complex and is hard to interpret.

Image Source: https://builtin.com/machine-learning/deep-learning
Neural Networks
Neural networks are called neural because they imitate the neurons of the human brain. The neural network has many nodes, with each performing its own activation function. The nodes connect the different layers of the neural network, which consists of the input layer, hidden layers and output layer.
The neurons of the given layer perform the same activation function. The activation function of each neuron is nonlinear, contributing to the complex patterns of the Deep neural network.

Image Source: https://doc.comsol.com/6.3/doc/com.comsol.help.comsol/comsol_ref_definitions.21.057.html
How it works
Neural networks were introduced decades ago with very few layers of neurons. Today with increased technical power, large number of neuron layers can be handled with many hidden layers, which contributes to the name ‘deep’.
The input layer is the starting point, and it receives raw data and corresponds to the nodes in the input layer. The values of these nodes are passed through the network for processing. The number of neurons in the input layer depends on the input features.
The hidden layers lie between the input layer and the output layer. They are responsible for transforming raw data into more meaningful and abstract data. Each neuron from the hidden layer is connected to each neuron in previous and succeeding hidden layer. The number of neurons is not required to be the same in the hidden layers.
The output layer is the final layer. It provides the desired output from the information passed through the hidden layers. The number of neurons in the output depends on the specific task.

Image Source: https://lamarr-institute.org/blog/deep-neural-networks/
Types of Neural Networks Models
· Feedforward neural networks (FNNs): They are the simplest type, where data flows in one direction from input to output. It is used for basic tasks like classification.
· Convolutional Neural Networks (CNNs): They can recognize pattern data and are ideal for computer vision. They are used for recognizing, processing and analyzing images. They are such a powerful tool that requires millions of labelled data for training.

· Recurrent Neural Networks (RNNs): They are used for processing sequential data, such as time series, speech recognition and natural language processing. RNNs have loops, such as the output of the given step serves as the input of the following computation step, to retain information over time. Variants like LSTMs and GRUs address vanishing gradient issues.

Image Source: https://botpenguin.com/glossary/recurrent-neural-network
· Autoencoders: They are used to compress the input data and reconstruct the original input using compressed recognition. These are applied in data compression, dimensionality reduction, feature extraction and fraud detection.

Image Source: https://www.grammarly.com/blog/ai/what-is-autoencoder/
· Generative Adversarial Networks (GANs): This network is used to create new data that resembles the original data. It consists of two networks—a generator (creates new data point as the original image and interact with the discriminator) and a discriminator (generator is frozen and the two images – original and the generated output- are provided and the network is trained to identify the fake image). GANs are widely used for image generation, style transfer and data augmentation.

Image Source: https://www.nature.com/articles/s41524-020-00352-0
· Transformer Networks: They are associated with Large Language Models (LLM) and have revolutionized NLP with self-attention mechanisms. This attention mechanism enables it to focus on a part of the input data that is more relevant to the given moment. Transformers excel at tasks like translation, text generation and sentiment analysis, powering models like GPT and BERT.
How It Is Trained
The connections between the nodes are represented by weights which are adjusted during the training to optimize the model’s performance. The neuron is a graphical representation of a numerical value. The weights that are the connection between the neurons are also just numerical values. The set of weights are different for every task and data sets.
The weights of the neural connections can change and so does the strength of the connections. The goal of training a model is to reduce errors and increase optimization. But to think training a model with such large capacity is overwhelming. This challenge was accomplished by using large number of labelled data and optimizing algorithms like backpropagation and gradient descent.
Backpropagation is a method to calculate how changing the individual weight in the neural network affects the accuracy of the model. It entails single end to end backward pass from the output of the loss function working back to the input layer. It describes how increasing or decreasing the neurons activation function in the output affects the overall loss.
Gradient Descent is a method in which the gradient calculated from the backpropagation is used as an input for this algorithm. It works by descending the gradient of the loss function which in turn reduces the error. The goal is to update the weight until you find the minimum gradient and find the parameter adjustments that contribute efficiently to minimum gradient.
Advantages
The high accuracy and optimized output in various tasks
Automatic feature learning from data eliminating manual intervention
Ability to handle large and complex data
Handling both structured and unstructured data for a large variety of tasks.
Disadvantages
It requires large data with high quality to be trained.
It is computationally expensive because it requires specialized hardware like GPUs and TPUs for training
Deep learning models are complex and hard to interpret results, it works like a black box.
When the model is trained repeatedly it becomes too specialized for the training data leading to poor performance on new data.
It relies on large data leading to concern over data privacy.
Domain expertise is required for successful implementation of deep learning.
References:


