A deep Q-learning neural network for Gymnasium's blackjack environment
Model: 4 linear layer model
Loss: Mean squared error
Trained on 1000 games then visually plays 1000 games
Note: Here instead of having a dedicated target and current network I just have one for simplicity.