You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Hi — has anyone run into repeated NaN losses when training mean_embedding_forecast_lstm with head: cmal on the full MultiMet Caravan set (6804 basins)? Config is close to floodhub-settings-config.yml with hidden_size: 512, batch_size: 512, max_updates_per_epoch: 2000, clip_gradient_norm: 1,validate_n_random_basins: -1 What I found with detect_anomaly: True: the traceback pointed to torch.log(1.0 - t) in training/loss.py. In modelzoo/head.py, tau is constructed as t = (1 - self._eps) * torch.sigmoid(t_latent) + self._eps which floors tau at _eps but allows it to reach exactly 1.0 when the sigmoid saturates, making log(1 - t) infinite. Is that the intended behaviour? Changing it to (1 - 2 * self._eps) * sigmoid(...) + self._eps bounds both ends. Is that the intended behaviour? What I've tried: raising _eps from 1e-5 to 1e-4; the symmetric tau bound above; clamping log_like at −50; reducing initial_learning_rate from 0.0005 to 0.0001. Training survives longer but still eventually fails. [ERROR] 22:38:57.605 (logging_utils.py:exception_logging) -- Uncaught exception
Traceback (most recent call last):
File "/usr/local/bin/run", line 7, in
sys.exit(_main())
^^^^^^^
File "/flood-forecasting/googlehydrology/run.py", line 122, in _main
continue_run(
File "/flood-forecasting/googlehydrology/run.py", line 196, in continue_run
start_training(base_config)
File "/flood-forecasting/googlehydrology/training/train.py", line 34, in start_training
trainer.train_and_validate()
File "/flood-forecasting/googlehydrology/training/basetrainer.py", line 325, in train_and_validate
self._train_epoch(epoch=epoch)
File "/flood-forecasting/googlehydrology/training/basetrainer.py", line 469, in _train_epoch
raise RuntimeError(
RuntimeError: Loss was NaN for 1 times in a row. Stopped training.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Hi — has anyone run into repeated NaN losses when training mean_embedding_forecast_lstm with head: cmal on the full MultiMet Caravan set (6804 basins)? Config is close to floodhub-settings-config.yml with hidden_size: 512, batch_size: 512, max_updates_per_epoch: 2000, clip_gradient_norm: 1,validate_n_random_basins: -1 What I found with detect_anomaly: True: the traceback pointed to torch.log(1.0 - t) in training/loss.py. In modelzoo/head.py, tau is constructed as t = (1 - self._eps) * torch.sigmoid(t_latent) + self._eps which floors tau at _eps but allows it to reach exactly 1.0 when the sigmoid saturates, making log(1 - t) infinite. Is that the intended behaviour? Changing it to (1 - 2 * self._eps) * sigmoid(...) + self._eps bounds both ends. Is that the intended behaviour? What I've tried: raising _eps from 1e-5 to 1e-4; the symmetric tau bound above; clamping log_like at −50; reducing initial_learning_rate from 0.0005 to 0.0001. Training survives longer but still eventually fails. [ERROR] 22:38:57.605 (logging_utils.py:exception_logging) -- Uncaught exception
Traceback (most recent call last):
File "/usr/local/bin/run", line 7, in
sys.exit(_main())
^^^^^^^
File "/flood-forecasting/googlehydrology/run.py", line 122, in _main
continue_run(
File "/flood-forecasting/googlehydrology/run.py", line 196, in continue_run
start_training(base_config)
File "/flood-forecasting/googlehydrology/training/train.py", line 34, in start_training
trainer.train_and_validate()
File "/flood-forecasting/googlehydrology/training/basetrainer.py", line 325, in train_and_validate
self._train_epoch(epoch=epoch)
File "/flood-forecasting/googlehydrology/training/basetrainer.py", line 469, in _train_epoch
raise RuntimeError(
RuntimeError: Loss was NaN for 1 times in a row. Stopped training.
All reactions