Neural Networks (and more!) with PyTorchRedes Neurais (e mais!) com PyTorch
Short Course - XI WPSMMinicurso - XI WPSM
Victor Coscrato
Introduction: What is PyTorch?Introdução: O que é PyTorch?
What is PyTorch?O que é PyTorch?
A high-performance numerical computing library for Python.Uma biblioteca de computação numérica de alto desempenho para Python.
Origins: Created by Facebook’s AI Research lab (FAIR) to overcome the flexibility limitations of existing frameworks at the time.Origens: Criado pelo laboratório de pesquisa em IA do Facebook (FAIR) para superar as limitações de flexibilidade de frameworks da época.
Open Source: It has been an open-source library since its inception in 2016, enabling rapid adoption and contributions from the global community.Open Source: É uma biblioteca de código aberto desde o seu nascimento em 2016, o que permitiu uma rápida adoção e contribuição da comunidade global.
Governance: Although it has always been open, governance shifted in 2022 from Meta to the PyTorch Foundation (under the Linux Foundation), ensuring neutral and collaborative management across multiple companies.Governança: Embora sempre tenha sido aberto, a governança mudou em 2022 da Meta para a PyTorch Foundation (sob a Linux Foundation), garantindo uma gestão neutra e colaborativa entre diversas empresas.
Key differentiator: Introduced the concept of dynamic computation graphs.Diferencial: Introduziu o conceito de grafos de computação dinâmicos.
Why PyTorch?Por que PyTorch?
Automatic differentiation (autograd): can automatically compute the gradient (the derivative) of any function you define. This eliminates the need to manually derive complex loss functions, a tedious and extremely error-prone process.Diferenciação automática (autograd): pode calcular automaticamente o gradiente (a derivada) de qualquer função que você definir. Isso elimina a necessidade de derivar manualmente funções de perda complexas, um processo tedioso e extremamente propenso a erros.
Hardware acceleration (GPU): enables the execution of mathematical operations on tensors (multidimensional arrays, like NumPy’s) in parallel on GPUs, which is essential for training deep neural network models in a reasonable time.Aceleração por hardware (GPU): permite a execução de operações matemáticas em tensores (arrays multidimensionais, como os do NumPy) de forma paralela em GPUs, o que é essencial para treinar modelos de redes neurais profundas em um tempo razoável.
Flexibility: It is hard to think of a parametric model that cannot be implemented in PyTorch.Flexibilidade: É difícil pensar em um modelo paramétrico que não possa ser implementado em PyTorch.
Tip
Think of PyTorch as NumPy with “superpowers”.Vale pensar no PyTorch como um NumPy com “superpoderes”.
Why PyTorch?Por que PyTorch?
Maturity: mature and stable, with a large and active community.Maturidade: madura e estável, com uma comunidade grande e ativa.
Popularity: dominates academic research, e.g. 80% of NeurIPS 2023 papers use PyTorch.Popularidade: domina a pesquisa acadêmica, e.g. 80% dos artigos do NeurIPS 2023 usam PyTorch.
Momentum and feedback loop: maturity and popularity create a virtuous cycle of development and adoption.Momentum e feedback loop: a maturidade e popularidade criam um ciclo virtuoso de desenvolvimento e adoção.
Source: demo video fromFonte: vídeo de demonstração daUltralytics.
Examples - Generative AIExemplos - IA Generativa
OK, but why should I care?Tá, mas e eu com isso?
Tip
Even if you don’t plan to train deep neural networks, PyTorch can help you – a lot.Mesmo que você não pretenda treinar redes neurais profundas, o PyTorch pode te ajudar, e muito.
Important
This short course will cover not only neural networks, but also everyday probabilistic and statistical applications.Esse minicurso abordará não apenas redes neurais, mas também aplicações probabilísticas e estatísticas cotidianas.
Part 1: PyTorch BasicsParte 1: Conceitos Básicos de PyTorch
TensorsTensores
The fundamental data unit in PyTorch. Analogous to NumPy’s ndarrays, but with “superpowers”.A unidade de dados no PyTorch. Análogo aos ndarrays do NumPy, mas com “superpoderes”.
A tensor has essential attributes:Um tensor possui atributos essenciais:
shape: the dimensionality of the tensor (e.g., scalar, vector, matrix).shape: a dimensionalidade do tensor (e.g., escalar, vetor, matriz).
dtype: the data type (torch.float32, torch.long, etc.).dtype: o tipo de dados (torch.float32, torch.long, etc.).
device: where the tensor is allocated in memory (cpu or cuda).device: onde o tensor está alocado na memória (cpu ou cuda).
TensorsTensores
# Creating tensors in various waysx = torch.tensor([[1, 2], [3, 4]], dtype=torch.float32)print(f"Tensor from list:\n{x}\n")# Attributesprint(f"Shape: {x.shape}")print(f"Data type: {x.dtype}")print(f"Device: {x.device}\n")# Interoperability with NumPy (they share the same memory on CPU)a_np = np.array([5, 6, 7])a_pt = torch.from_numpy(a_np)print(f"Tensor from NumPy: {a_pt}")a_pt[0] =99print(f"Original NumPy was modified: {a_np}")
Tensor from list:
tensor([[1., 2.],
[3., 4.]])
Shape: torch.Size([2, 2])
Data type: torch.float32
Device: cpu
Tensor from NumPy: tensor([5, 6, 7])
Original NumPy was modified: [99 6 7]
Tensor OperationsOperações com Tensores
The syntax is familiar to NumPy users.A sintaxe é familiar para quem usa NumPy.
x = torch.tensor([[1., 2.], [3., 4.]])y = torch.tensor([[5., 6.], [7., 8.]])# Element-wise operationsprint("Sum:", x + y)print("Product (Hadamard):", x * y)# Matrix multiplicationprint("Matrix Product (@):", x @ y)print("Matrix Product (matmul):", torch.matmul(x, y))# Reshapingprint("Original shape:", x.shape)print("Reshaped (view):", x.view(4, 1).shape)
Moving tensors to the GPU is trivial and can result in speed gains of orders of magnitude.Mover tensores para a GPU é trivial e pode resultar em ganhos de velocidade de ordens de magnitude.
# Check if GPU is availabledevice ="cuda"if torch.cuda.is_available() else"cpu"print(f"Using device: {device}")# Create tensors on the chosen devicex_gpu = torch.randn(1000, 1000, device=device)y_gpu = torch.randn(1000, 1000, device=device)# The operation is executed on the devicez_gpu = x_gpu @ y_gpu# To use with libraries like Matplotlib or NumPy, move back to CPUz_cpu = z_gpu.to("cpu")print(z_cpu.shape)
Using device: cuda
torch.Size([1000, 1000])
The Tensor “Superpower”O “superpoder” dos tensores
Important
autograd is the system that tracks all operations on tensors to automatically compute gradients. It is the foundation for model optimization.autograd é o sistema que rastreia todas as operações em tensores para calcular gradientes automaticamente. É a base para a otimização de modelos.
Dynamic computation graph: PyTorch builds a graph “on-the-fly”.Grafo computacional dinâmico: PyTorch constrói um grafo “on-the-fly”.
Tracking: for a tensor x, if x.requires_grad=True, PyTorch records its operation “history”.Rastreamento: para um tensor x, se x.requires_grad=True, PyTorch guarda sua “história” de operações.
Gradient computation: by calling .backward() on a scalar (e.g., loss function), PyTorch applies the chain rule.Cálculo de gradientes: ao chamar .backward() em um escalar (e.g., função de perda), PyTorch aplica a regra da cadeia.
autograd in Practiceautograd na Prática
Let’s compute the derivative of \(y = x^2 \sin(x)\) at \(x = 2\).Vamos calcular a derivada de \(y = x^2 \sin(x)\) em \(x = 2\).
Tip
The analytical derivative is \(2x \sin(x) + x^2 \cos(x)\).A derivada analítica é \(2x \sin(x) + x^2 \cos(x)\).
# Define the tensor and enable gradient trackingx = torch.tensor(2.0, requires_grad=True)# Define the functiony = x**2* torch.sin(x)# Compute the gradienty.backward()# The gradient is stored in x.gradprint(f"Gradient of y with respect to x at x=2: {x.grad.item():.4f}")# Analytical value for comparisonanalytic_grad =2*2* np.sin(2) +2**2* np.cos(2)print(f"Analytical value: {analytic_grad:.4f}")
Gradient of y with respect to x at x=2: 1.9726
Analytical value: 1.9726
Best Practices with TensorsBoas Práticas com Tensores
requires_grad=True only for parameters: input data generally does not need gradients.requires_grad=True apenas para parâmetros: dados de entrada geralmente não precisam de gradiente.
Use with torch.no_grad() during inference to save memory and time.Use with torch.no_grad() em inferência para economizar memória e tempo.
w = torch.tensor(1.0, requires_grad=True)x_obs = torch.tensor(2.0) # observed datay_hat = w * x_obsy_hat.backward() # We'll see more about this laterprint("gradient at w:", w.grad)with torch.no_grad(): pred = w *10print("prediction without tracking gradient:", pred)
gradient at w: tensor(2.)
prediction without tracking gradient: tensor(10.)
The Dynamic Graph and grad_fnO Grafo Dinâmico e o grad_fn
In PyTorch, results of operations on tensors with gradients carry a pointer to the operation that created them.No PyTorch, resultados de operações sobre tensores com gradiente carregam um ponteiro para a operação que os criou.
a = torch.tensor(5.0, requires_grad=True)b = torch.tensor(3.0, requires_grad=True)d = a * b + aprint("d:", d)print("grad_fn of d:", d.grad_fn)
d: tensor(20., grad_fn=<AddBackward0>)
grad_fn of d: <AddBackward0 object at 0x7f977c6a48e0>
Visualizing the Computation GraphVisualizando o Grafo Computacional
make_dot(d, params={'d': d, 'a': a, 'b': b})
Blue rectangles are the leaves (our parameters)Os retângulos azuis são as folhas (nossos parâmetros)
Gray rectangles are the intermediate operationsOs retângulos cinzas são as operações intermediárias
The arrow indicates the information flow (forward pass)A seta indica o fluxo de informação (forward pass)
Computing the DerivativesCalculando as derivadas
make_dot(d, params={'d': d, 'a': a, 'b': b})
We can compute gradients automatically through the backward pass: PyTorch traverses the graph in reverse and applies the chain rule to compute partial derivatives.Podemos calcular gradientes automaticamente através do backward pass: o PyTorch percorre o grafo de trás para frente e aplica a regra da cadeia para calcular as derivadas parciais.
Tensors are PyTorch’s core data structure (shape, dtype, device).Tensores são a estrutura central do PyTorch (shape, dtype, device).
requires_grad=True enables tracking for parameters that will be optimized.requires_grad=True ativa o rastreamento para parâmetros que serão otimizados.
autograd builds the dynamic graph and computes gradients with .backward().O autograd constrói o grafo dinâmico e calcula gradientes com .backward().
Important
With this, we already have the basics to define models and train parameters via numerical optimization.Com isso, já temos o básico para definir modelos e treinar parâmetros por otimização numérica.
Part 2: Statistical Models with PyTorchParte 2: Modelos Estatísticos com PyTorch
Example: Linear RegressionExemplo: Regressão Linear
Let’s recall the analytical least squares solution for the linear regression model:Vamos recordar a solução analítica de mínimos quadrados do modelo de regressão linear:
\[
y = X \beta + \epsilon
\]
The quadratic loss is:A perda quadrática é:\[
L(\beta) = \sum_{i=1}^{n} (y_i - X_i^T\beta)^2
\]
Differentiating and setting to zero, we obtain the closed-form solution:Derivando e igualando a zero, obtemos a solução fechada:
\[
\hat{\beta} = (X^T X)^{-1}X^Ty
\]
Let’s compare this solution with optimization via autograd.Vamos comparar essa solução com a otimização via autograd.
Optimized beta via autograd:
[[ 1.9843516]
[-3.5014331]
[ 1.0013602]]
Interpretation and Practical GainInterpretação e Ganho Prático
In the example, autograd recovers a solution very close to the analytical one.No exemplo, o autograd recupera uma solução muito próxima da analítica.
The real gain is not finding the least squares solution, but having a universal optimizer for models without closed-form solutions.O ganho real não é encontrar a solução de mínimos quadrados, mas ter um otimizador universal para modelos sem solução fechada.
The same code pattern extends to logistic regression, neural networks, and custom models.O mesmo padrão de código se estende para regressão logística, redes neurais e modelos customizados.
Optimized beta via autograd:
[[-1.9755253]
[ 1.5227022]
[-0.9161041]] [-0.06488772]
Transition to torch.nnTransição para torch.nn
So far, we have worked with tensors and autograd directly. The torch.nn API organizes this into reusable modules. Let’s repeat logistic regression using nn.Module.Até aqui, trabalhamos com tensores e autograd de forma direta. A API torch.nn organiza isso em módulos reutilizáveis. Vamos repetir a regressão logística usando nn.Module.
class LogisticRegressionModel(nn.Module):def__init__(self, n_features):super(LogisticRegressionModel, self).__init__()self.linear = nn.Linear(n_features, 1)def forward(self, x):return torch.sigmoid(self.linear(x))
Logistic Regression with torch.nnRegressão Logística com torch.nn
model_logistic = LogisticRegressionModel(n_features)criterion = nn.BCELoss()optimizer = optim.SGD(model_logistic.parameters(), lr=0.1)n_epochs =10000for epoch inrange(n_epochs): optimizer.zero_grad() y_pred = model_logistic(X_tensor) loss = criterion(y_pred, y_tensor) loss.backward() optimizer.step()print("Optimized beta via torch.nn for logistic regression:")print(model_logistic.linear.weight.detach().numpy(), model_logistic.linear.bias.detach().numpy())# make_dot(loss, params=dict(model_logistic.named_parameters()))
Optimized beta via torch.nn for logistic regression:
[[-1.9755336 1.5226941 -0.91611785]] [-0.06487162]
The torch.nn APIA API torch.nn
nn.Module: standard container for parameters and forward logic.nn.Module: container padrão para parâmetros e lógica forward.
nn.Linear: pre-built module with weights and bias.nn.Linear: módulo pré-construído com pesos e intercepto.
torch.optim: parameter updates + zero_grad().torch.optim: atualização dos parâmetros + zero_grad().
Part 3: Neural Networks with PyTorchParte 3: Redes Neurais com PyTorch
Review: Neural NetworksRevisão: Redes Neurais
A neuron receives inputs \(\mathbf{x} \in \mathbb{R}^d\), computes a parametric linear combination, and applies a non-linear transformation \(\phi\):Um neurônio recebe entradas \(\mathbf{x} \in \mathbb{R}^d\), calcula uma combinação linear paramétrica e aplica uma transformação não-linear \(\phi\):\[ y = \phi(\mathbf{w}^T\mathbf{x} + b) \]
The function \(\phi\) (e.g., ELU, ReLU, \(\tanh\)) is essential. Without it, a composition of neurons would collapse into a single affine function:A função \(\phi\) (e.g., ELU, ReLU, \(\tanh\)) é essencial. Sem ela, uma composição de neurônios colapsaria em uma única função afim:\[ (W_2 (W_1 \mathbf{x} + b_1) + b_2) = (W_2 W_1)\mathbf{x} + (W_2 b_1 + b_2) = W'\mathbf{x} + b' \]
Note
Logistic regression is essentially a single neuron with \(\phi = \sigma\) (sigmoid function), where:A regressão logística é essencialmente um único neurônio com \(\phi = \sigma\) (função sigmóide), onde:\[ P(y=1|\mathbf{x}) = \sigma(\beta^T\mathbf{x}) \]
Review: Neural NetworksRevisão: Redes Neurais
Universal approximation theorem: A feedforward neural network with at least one hidden layer and a non-linear activation function is capable of approximating any continuous function \(\mathbb{R}^n \to \mathbb{R}^m\), given enough neurons.Teorema da aproximação universal: Uma rede neural do tipo feedforward com pelo menos uma camada oculta e uma função de ativação não-linear é capaz de aproximar qualquer função contínua \(\mathbb{R}^n \to \mathbb{R}^m\), desde que haja neurônios o suficiente.
Deep networks: In practice, multiple layers tend to be a more efficient representation, with better generalization.Redes profundas: Na prática, múltiplas camadas tendem a ser uma representação mais eficiente, e com melhor generalização.
Multi-Layer Perceptron
An MLP with two hidden layers maps the data \(\mathbf{X}\) through the operations:Um MLP com duas camadas ocultas mapeia os dados \(\mathbf{X}\) através das operações:
How do we optimize the parameters \(\mathbf{W}_1, \mathbf{W}_2, \mathbf{W}_3\) of a neural network? We need the gradients \(\nabla_{\mathbf{W}_k} \ell\) for each layer.Como otimizar os parâmetros \(\mathbf{W}_1, \mathbf{W}_2, \mathbf{W}_3\) de uma rede neural? Precisamos dos gradientes \(\nabla_{\mathbf{W}_k} \ell\) para cada camada.
Consider the quadratic loss as an example:Considere a perda quadrática como exemplo:
Substituting \(\hat{\mathbf{y}}\) with the MLP expression, the loss becomes a composition of functions:Substituindo \(\hat{\mathbf{y}}\) pela expressão do MLP, a perda se torna uma composição de funções:
By the chain rule, gradients propagate layer by layer, from back to front:Pela regra da cadeia, os gradientes se propagam camada a camada, de trás para frente:
This is exactly the computation that autograd performs when we call loss.backward(). No matter how many layers the network has, the chain rule applies recursively.Esse é exatamente o cálculo que o autograd realiza ao chamarmos loss.backward(). Não importa quantas camadas a rede tenha, a regra da cadeia se aplica recursivamente.
OptimizersOtimizadores
With the gradients \(\nabla_\theta \ell\) in hand, we need an update rule for the parameters.Com os gradientes \(\nabla_\theta \ell\) em mãos, precisamos de uma regra de atualização para os parâmetros.
SGD (Stochastic Gradient Descent): the simplest update:SGD (Stochastic Gradient Descent): a atualização mais simples:\[ \theta \leftarrow \theta - \eta \, \nabla_\theta \ell \]where \(\eta\) is the learning rate.onde \(\eta\) é a taxa de aprendizado (learning rate).
Adam: combines moving averages of the gradient (\(1^{\text{st}}\) moment) and the squared gradient (\(2^{\text{nd}}\) moment), adapting the learning rate for each parameter individually. It is the default optimizer in most modern applications.Adam: combina médias móveis do gradiente (\(1^{\circ}\) momento) e do gradiente ao quadrado (\(2^{\circ}\) momento), adaptando a taxa de aprendizado para cada parâmetro individualmente. É o otimizador padrão na maioria das aplicações modernas.
In practice, just swap optim.SGD(...) for optim.Adam(...), the API is identical.Na prática, basta trocar optim.SGD(...) por optim.Adam(...), a API é idêntica.
Important
There is no need to implement backpropagation or optimizers manually: PyTorch handles everything. Just define the network, the loss, and call loss.backward() + optimizer.step().Não é necessário implementar backpropagation nem os otimizadores manualmente: o PyTorch cuida de tudo. Basta definir a rede, a perda, e chamar loss.backward() + optimizer.step().
MLP for Binary ClassificationMLP para Classificação Binária
MLP for Binary ClassificationMLP para Classificação Binária
Recall…Relembrando…
autograd automatically differentiates through layers and activations. Just define the loss and call loss.backward().O autograd diferencia automaticamente através de camadas e ativações. Basta definir a perda e chamar loss.backward().
Exercises: Optimization with PyTorchExercícios: Otimização com PyTorch
1: Poisson Regression (GLM)1: Regressão de Poisson (GLM)
Objective: implement and optimize a Poisson regression with a custom loss function.Objetivo: implementar e otimizar uma regressão de Poisson com função de perda customizada.
Train with gradient descent and visualize convergence.Treinar com gradiente descendente e visualizar convergência.
2: Gaussian Mixture Model (GMM)
Objective: fit the parameters of a 1D GMM with two components directly via autograd.Objetivo: ajustar os parâmetros de um GMM 1D com duas componentes diretamente via autograd.
To enforce constraints:Para impor restrições: - softmax for mixture weights.softmax para pesos da mistura. - exp on log-standard deviations to ensure \(\sigma_k > 0\).exp nos log-desvios para garantir \(\sigma_k > 0\).
TaskTarefa
Generate data from a 1D GMM with two Gaussians.Gerar dados de um GMM 1D com duas gaussianas.
Define parameters as torch.nn.Parameter.Definir parâmetros como torch.nn.Parameter.
Implement gmm_nll_loss with torch.logsumexp.Implementar gmm_nll_loss com torch.logsumexp.
Optimize parameters and compare learned density with actual data.Otimizar parâmetros e comparar densidade aprendida com dados reais.
3: Heteroscedastic Linear Regression3: Regressão Linear Heterocedástica
Objective: model conditional mean and variance simultaneously.Objetivo: modelar média e variância condicional simultaneamente.
Train on simulated data with increasing variance.Treinar em dados simulados com variância crescente.
Plot predicted mean and uncertainty bands (\(\pm 2\sigma\)).Plotar média prevista e bandas de incerteza (\(\pm 2\sigma\)).
4: CNN (MNIST)
Objective: build and train a CNN to classify digits.Objetivo: construir e treinar uma CNN para classificar dígitos.
Given an image \(\mathbf{X} \in \mathbb{R}^{28 \times 28}\), classify it (0 to 9) by minimizing the Cross-Entropy Loss:Dada uma imagem \(\mathbf{X} \in \mathbb{R}^{28 \times 28}\), classifique-a (0 a 9) minimizando a entropia cruzada (Cross-Entropy Loss):\[ L = -\frac{1}{n}\sum_{i,c} y_{i,c} \log(\hat{y}_{i,c}) \]
TaskTarefa
Load MNIST via torchvision.datasets.Carregar o MNIST via torchvision.datasets.
Create a model with nn.Conv2d, nn.MaxPool2d, and nn.Linear.Criar modelo com nn.Conv2d, nn.MaxPool2d e nn.Linear.
Train in mini-batches (DataLoader) with nn.CrossEntropyLoss().Treinar em mini-batches (DataLoader) com nn.CrossEntropyLoss().
Evaluate accuracy on the test set.Avaliar a acurácia no teste.
5: MF for Movie Recommendation5: FM para Recomendação de filmes
Objective: apply matrix factorization for recommendation systems.Objetivo: aplicar fatorização matricial para sistemas de recomendação.
We approximate the rating \(r_{u,i}\) (user \(u\), item \(i\)) by:Aproximamos a avaliação \(r_{u,i}\) (usuário \(u\), item \(i\)) por:\[ \hat{r}_{u,i} = \mu + b_u + b_i + \mathbf{p}_u^\top \mathbf{q}_i \]
Minimizing the mean squared error with \(L_2\) regularization:Minimizando o erro quadrático médio com penalização (\(L_2\)):\[ L = \sum_{(u,i)} (r_{u,i} - \hat{r}_{u,i})^2 + \lambda \left(||\mathbf{p}_u||^2 + ||\mathbf{q}_i||^2 + b_u^2 + b_i^2\right) \]
TaskTarefa
Use nn.Embedding for \(p_u, q_i, b_u, b_i\).Usar nn.Embedding para \(p_u, q_i, b_u, b_i\).
Compute \(\hat{r}_{u,i}\) in forward and optimize the loss with MSE.Calcular \(\hat{r}_{u,i}\) no forward e otimizar a perda com MSE.
Train on the movielens-100k dataset.Treinar no conjunto movielens-100k.
Provide top-K recommendations for a user.Indicar top-K recomendações para um usuário.
6: RNN for Time Series6: RNN para Séries Temporais
Objective: predict values in a time series using recurrent networks.Objetivo: prever valores em uma série temporal usando redes recorrentes.
Given a sequence \((x_{t-k}, \dots, x_{t})\), predict \(x_{t+1}\) by updating the hidden state \(\mathbf{h}_t\):Dada uma sequência \((x_{t-k}, \dots, x_{t})\), preveja \(x_{t+1}\) atualizando o estado oculto \(\mathbf{h}_t\):\[ \mathbf{h}_t = f(\mathbf{x}_t, \mathbf{h}_{t-1}), \quad \hat{x}_{t+1} = \mathbf{w}^\top \mathbf{h}_t + b \]
TaskTarefa
Use the Sunspots dataset (statsmodels.api.datasets.sunspots).Utilizar os dados de Manchas Solares (statsmodels.api.datasets.sunspots).
Format data using history windows.Formatar dados usando janelas do histórico.
Implement a model with nn.RNN or nn.LSTM and nn.Linear.Implementar modelo com nn.RNN ou nn.LSTM e nn.Linear.
Train by minimizing MSE and plot predictions against original data.Treinar minimizando MSE e plotar previsões contra dados originais.