Yet-another-efficientdet-pytorch: unstable loss during training epoches

Created on 13 Apr 2020  路  5Comments  路  Source: zylo117/Yet-Another-EfficientDet-Pytorch

Hi @zylo117 ,

I found the training loss changes dynamically during training epoches.
e.g. efficient-d1, during last several batch iters of epoch 0, the loss about 0.3, 0.08
batch size 32, init lr 0.005

then during the epoch 1, it will be much larger than that, i have tried several times and this will happen every time. My custom data should be ok, since offical tf implement could be trained with same dataset normally.
Do you have any clues for this>

Step: 895. Epoch: 0/500. Iteration: 896/897. Cls loss: 0.31577. Reg loss: 0.06190. Total loss: 0.37767: 100%|鈻墊 Step: 895. Epoch: 0/500. Iteration: 896/897. Cls loss: 0.31577. Reg loss: 0.06190. Total loss: 0.37767: 100%|鈻墊 Step: 896. Epoch: 0/500. Iteration: 897/897. Cls loss: 0.34984. Reg loss: 0.07105. Total loss: 0.42089: 100%|鈻墊 Step: 896. Epoch: 0/500. Iteration: 897/897. Cls loss: 0.34984. Reg loss: 0.07105. Total loss: 0.42089: 100%|鈻坾 897/897 [17:16<00:00, 1.16s/it]
^@Val. Epoch: 0/500. Classification loss: 0.32317. Regression loss: 0.14466. Total loss: 0.46783
Step: 897. Epoch: 1/500. Iteration: 1/897. Cls loss: 0.34505. Reg loss: 0.16021. Total loss: 0.50526: 0%| | 0/Step: 897. Epoch: 1/500. Iteration: 1/897. Cls loss: 0.34505. Reg loss: 0.16021. Total loss: 0.50526: 0%| | 1/Step: 898. Epoch: 1/500. Iteration: 2/897. Cls loss: 1.12159. Reg loss: 2.12988. Total loss: 3.25147: 0%| | 1/Step: 898. Epoch: 1/500. Iteration: 2/897. Cls loss: 1.12159. Reg loss: 2.12988. Total loss: 3.25147: 0%| | 2/Step: 899. Epoch: 1/500. Iteration: 3/897. Cls loss: 2.10690. Reg loss: 0.88227. Total loss: 2.98917: 0%| | 2/Step: 899. Epoch: 1/500. Iteration: 3/897. Cls loss: 2.10690. Reg loss: 0.88227. Total loss: 2.98917: 0%| | 3/Step: 900. Epoch: 1/500. Iteration: 4/897. Cls loss: 2.30212. Reg loss: 80.03463. Total loss: 82.33675: 0%| | Step: 900. Epoch: 1/500. Iteration: 4/897. Cls loss: 2.30212. Reg loss: 80.03463. Total loss: 82.33675: 0%| | Step: 901. Epoch: 1/500. Iteration: 5/897. Cls loss: 2.30212. Reg loss: 67977.46094. Total loss: 67979.76562: Step: 901. Epoch: 1/500. Iteration: 5/897. Cls loss: 2.30212. Reg loss: 67977.46094. Total loss: 67979.76562: Step: 902. Epoch: 1/500. Iteration: 6/897. Cls loss: 2.30212. Reg loss: 65421.88281. Total loss: 65424.18359: Step: 902. Epoch: 1/500. Iteration: 6/897. Cls loss: 2.30212. Reg loss: 65421.88281. Total loss: 65424.18359: Step: 903. Epoch: 1/500. Iteration: 7/897. Cls loss: 2.30212. Reg loss: 1573417.00000. Total loss: 1573419.25000Step: 903. Epoch: 1/500. Iteration: 7/897. Cls loss: 2.30212. Reg loss: 1573417.00000. Total loss: 1573419.25000Step: 904. Epoch: 1/500. Iteration: 8/897. Cls loss: 2.70166. Reg loss: 272323072.00000. Total loss: 272323072.0Step: 904. Epoch: 1/500. Iteration: 8/897. Cls loss: 2.70166. Reg loss: 272323072.00000. Total loss: 272323072.0Step: 905. Epoch: 1/500. Iteration: 9/897. Cls loss: 2.78378. Reg loss: 5323778686976.00000. Total loss: 5323778

Step: 1142. Epoch: 1/500. Iteration: 246/897. Cls loss: 62.49104. Reg loss: 0.50637. Total loss: 62.99741: 27%|Step: 1143. Epoch: 1/500. Iteration: 247/897. Cls loss: 73.22089. Reg loss: 1213188314532348763832320.00000. TotStep: 1143. Epoch: 1/500. Iteration: 247/897. Cls loss: 73.22089. Reg loss: 1213188314532348763832320.00000. TotStep: 1144. Epoch: 1/500. Iteration: 248/897. Cls loss: 71.14322. Reg loss: 451076503452162755395584.00000. TotaStep: 1144. Epoch: 1/500. Iteration: 248/897. Cls loss: 71.14322. Reg loss: 451076503452162755395584.00000. TotaStep: 1145. Epoch: 1/500. Iteration: 249/897. Cls loss: 66.21103. Reg loss: 441545013143201799733248.00000. TotaStep: 1145. Epoch: 1/500. Iteration: 249/897. Cls loss: 66.21103. Reg loss: 441545013143201799733248.00000. TotaStep: 1146. Epoch: 1/500. Iteration: 250/897. Cls loss: 68.61477. Reg loss: 235899719417569696808960.00000. TotaStep: 1146. Epoch: 1/500. Iteration: 250/897. Cls loss: 68.61477. Reg loss: 235899719417569696808960.00000. TotaStep: 1147. Epoch: 1/500. Iteration: 251/897. Cls loss: 61.03159. Reg loss: 359504415912071251623936.00000. TotaStep: 1147. Epoch: 1/500. Iteration: 251/897. Cls loss: 61.03159. Reg loss: 359504415912071251623936.00000. TotaStep: 1148. Epoch: 1/500. Iteration: 252/897. Cls loss: 62.65581. Reg loss: 1484498270567262345756672.00000. TotStep: 1148. Epoch: 1/500. Iteration: 252/897. Cls loss: 62.65581. Reg loss: 1484498270567262345756672.00000. TotStep: 1149. Epoch: 1/500. Iteration: 253/897. Cls loss: 63.57681. Reg loss: 703274623820396710854656.00000. TotaStep: 1149. Epoch: 1/500. Iteration: 253/897. Cls loss: 63.57681. Reg loss: 703274623820396710854656.00000. TotaStep: 1150. Epoch: 1/500. Iteration: 254/897. Cls loss: 66.59061. Reg loss: 229218953629538736668672.00000. TotaStep: 1150. Epoch: 1/500. Iteration: 254/897. Cls loss: 66.59061. Reg loss: 229218953629538736668672.00000. TotaStep: 1151. Epoch: 1/500. Iteration: 255/897. Cls loss: 66.09634. Reg loss: 871639065787460319444992.00000. TotaStep: 1151. Epoch: 1/500. Iteration: 255/897. Cls loss: 66.09634. Reg loss: 871639065787460319444992.00000. TotaStep: 1152. Epoch: 1/500. Iteration: 256/897. Cls loss: 68.23049. Reg loss: 1114574254499712676659200.00000. TotStep: 1152. Epoch: 1/500. Iteration: 256/897. Cls loss: 68.23049. Reg loss: 1114574254499712676659200.00000. TotStep: 1153. Epoch: 1/500. Iteration: 257/897. Cls loss: 61.66665. Reg loss: 350685359035363289464832.00000. TotaStep: 1153. Epoch: 1/500. Iteration: 257/897. Cls loss: 61.66665. Reg loss: 350685359035363289464832.00000. Tot

Most helpful comment

sorry for the training troubles, there's a bug in loss function. please pull the latest code and give it a try.
and yes, smaller lr

Thanks very much for your reply, i will try the latest code.
You are the auther of this open source implement, you dont need to apologize.
Thanks again for your great work!

All 5 comments

reduce your learning rate by 100 times and try again

reduce your learning rate by 100 times and try again

Thanks...i use 0.001 and the loss seems like to be normal again.

Thanks very much for your help, great work!

Oh, sorry.
The loss gets larger(from 0.2(classificartion)+0.05(regression) to 2(classification)+1(regression)) after 13 epoches on custom dataset, and ap test result on val dataset does not change much.
Do i have to set smaller lr?

thanks very much for your help.

sorry for the training troubles, there's a bug in loss function. please pull the latest code and give it a try.
and yes, smaller lr

sorry for the training troubles, there's a bug in loss function. please pull the latest code and give it a try.
and yes, smaller lr

Thanks very much for your reply, i will try the latest code.
You are the auther of this open source implement, you dont need to apologize.
Thanks again for your great work!

Was this page helpful?
0 / 5 - 0 ratings

Related issues

guitar9 picture guitar9  路  11Comments

you-old picture you-old  路  4Comments

qtw1998 picture qtw1998  路  3Comments

chunleiml picture chunleiml  路  5Comments

lzneu picture lzneu  路  10Comments