[home]   [coding projects]   [research projects]


Multi-Layer Perceptrons

Activation functions, architecture shape, optimizers, and regularization on synthetic nonlinear tasks.

[full source]

This project builds a configurable multilayer perceptron in PyTorch and uses it as a controlled experiment platform. The data are synthetic so the architecture and optimization variables can be changed without dataset ambiguity dominating the comparison.

1. Synthetic nonlinear tasks

The classification labels combine a sinusoidal term with a linear term, while the regression target combines a sinusoid and a quadratic norm:

\[ y_{\mathrm{class}} = \mathbf 1\!\left[ \sin(w_1^\top x)+0.5\,w_2^\top x+\varepsilon>0 \right], \qquad y_{\mathrm{reg}} = \sin(w^\top x)+0.1\|x\|^2+\varepsilon. \]

Normalization statistics are computed from the training set and then reused for validation.

2. Configurable MLP

Christopher_Housholder_HW2.py
class MLP(nn.Module):
    def __init__(
        self,
        input_dim,
        hidden_layers,
        output_dim,
        activation="relu"
    ):
        super(MLP, self).__init__()

        match activation:
            case "relu": self.activation = nn.ReLU()
            case "sigmoid": self.activation = nn.Sigmoid()
            case "tanh": self.activation = nn.Tanh()

        layers = []
        prev_dim = input_dim

        for hidden_dim in hidden_layers:
            layers.append(nn.Linear(prev_dim, hidden_dim))
            layers.append(self.activation)
            prev_dim = hidden_dim

        layers.append(nn.Linear(prev_dim, output_dim))
        self.model = nn.Sequential(*layers)

Weight initialization follows the activation: He initialization for ReLU and Xavier initialization for sigmoid/tanh. The same training routine can switch between classification/regression losses, SGD/Adam, batch size, momentum, weight decay, and an optional convergence tolerance.

3. Experiment matrix

The experiments compare activation functions, shallow-wide versus deep-narrow networks, SGD versus Adam, and \(L_2\) regularization. The report's overall finding is that activation and optimization choice primarily affect training efficiency, while the tested depth difference produces a more modest representational gain on these synthetic tasks.

Activation-function comparison.
Activation-function comparison.
Depth-versus-width comparison.
Depth-versus-width comparison.
Optimizer comparison.
Optimizer comparison.
Regularization comparison.
Regularization comparison.

These pages are selective technical manuals: enough source to expose the mechanism, not a mirror of the entire repository.
Last updated: September 14, 2026.