[home] [coding projects] [research projects]
Activation functions, architecture shape, optimizers, and regularization on synthetic nonlinear tasks.
This project builds a configurable multilayer perceptron in PyTorch and uses it as a controlled experiment platform. The data are synthetic so the architecture and optimization variables can be changed without dataset ambiguity dominating the comparison.
The classification labels combine a sinusoidal term with a linear term, while the regression target combines a sinusoid and a quadratic norm:
Normalization statistics are computed from the training set and then reused for validation.
class MLP(nn.Module):
def __init__(
self,
input_dim,
hidden_layers,
output_dim,
activation="relu"
):
super(MLP, self).__init__()
match activation:
case "relu": self.activation = nn.ReLU()
case "sigmoid": self.activation = nn.Sigmoid()
case "tanh": self.activation = nn.Tanh()
layers = []
prev_dim = input_dim
for hidden_dim in hidden_layers:
layers.append(nn.Linear(prev_dim, hidden_dim))
layers.append(self.activation)
prev_dim = hidden_dim
layers.append(nn.Linear(prev_dim, output_dim))
self.model = nn.Sequential(*layers)
Weight initialization follows the activation: He initialization for ReLU and Xavier initialization for sigmoid/tanh. The same training routine can switch between classification/regression losses, SGD/Adam, batch size, momentum, weight decay, and an optional convergence tolerance.
The experiments compare activation functions, shallow-wide versus deep-narrow networks, SGD versus Adam, and \(L_2\) regularization. The report's overall finding is that activation and optimization choice primarily affect training efficiency, while the tested depth difference produces a more modest representational gain on these synthetic tasks.




These pages are selective technical manuals: enough source to expose the mechanism, not a mirror of the entire repository.
Last updated: September 14, 2026.