bigscience-bot commited on
Commit
36f6976
1 Parent(s): c101aef
Files changed (1) hide show
  1. logs/main_log.txt +12 -0
logs/main_log.txt CHANGED
@@ -66690,3 +66690,15 @@ time (ms)
66690
  time (ms)
66691
  [2021-09-25 18:51:17] PULSE: tr8-104B is scheduled to start in 19:55:07 (at 2021-09-26T14:46:25) (1188168 on 'gpu_p13' partition)
66692
  [2021-09-25 18:51:17] PULSE: tr8-104B is running for 14:24:16 since 2021-09-25T04:27:01 (1185639_1 on 'gpu_p13' partition (r6i5n[7-8],r6i6n0,r7i2n[4-5],r7i3n2,r7i6n[2-4],r7i7n[7-8],r8i0n[0,2-3,5-8],r8i1n[0,2-4],r8i2n8,r8i3n[0-2],r8i5n[3-4],r8i7n[3-8],r9i0n[0-5],r9i1n[0-3],r9i2n[3-6,8],r9i3n[0-1,7-8],r9i4n[0-3],r9i5n[3-8],r9i6n0)
 
 
 
 
 
 
 
 
 
 
 
 
 
66690
  time (ms)
66691
  [2021-09-25 18:51:17] PULSE: tr8-104B is scheduled to start in 19:55:07 (at 2021-09-26T14:46:25) (1188168 on 'gpu_p13' partition)
66692
  [2021-09-25 18:51:17] PULSE: tr8-104B is running for 14:24:16 since 2021-09-25T04:27:01 (1185639_1 on 'gpu_p13' partition (r6i5n[7-8],r6i6n0,r7i2n[4-5],r7i3n2,r7i6n[2-4],r7i7n[7-8],r8i0n[0,2-3,5-8],r8i1n[0,2-4],r8i2n8,r8i3n[0-2],r8i5n[3-4],r8i7n[3-8],r9i0n[0-5],r9i1n[0-3],r9i2n[3-6,8],r9i3n[0-1,7-8],r9i4n[0-3],r9i5n[3-8],r9i6n0)
66693
+ iteration 8800/ 159576 | consumed samples: 498960 | elapsed time per iteration (ms): 23872.5 | learning rate: 6.000E-05 | global batch size: 176 | lm loss: 7.150304E+00 | loss scale: 1024.0 | grad norm: 20582.002 | num zeros: 0.0 | number of skipped iterations: 0 | number of nan iterations: 0 |
66694
+ time (ms)
66695
+ iteration 8810/ 159576 | consumed samples: 500720 | elapsed time per iteration (ms): 23674.3 | learning rate: 6.000E-05 | global batch size: 176 | lm loss: 7.121466E+00 | loss scale: 1024.0 | grad norm: 26026.638 | num zeros: 0.0 | number of skipped iterations: 0 | number of nan iterations: 0 |
66696
+ time (ms)
66697
+ iteration 8820/ 159576 | consumed samples: 502480 | elapsed time per iteration (ms): 23655.3 | learning rate: 6.000E-05 | global batch size: 176 | lm loss: 7.227619E+00 | loss scale: 1024.0 | grad norm: 19493.231 | num zeros: 0.0 | number of skipped iterations: 0 | number of nan iterations: 0 |
66698
+ time (ms)
66699
+ iteration 8830/ 159576 | consumed samples: 504240 | elapsed time per iteration (ms): 24040.7 | learning rate: 6.000E-05 | global batch size: 176 | lm loss: 7.202127E+00 | loss scale: 1024.0 | grad norm: 21130.889 | num zeros: 0.0 | number of skipped iterations: 0 | number of nan iterations: 0 |
66700
+ time (ms)
66701
+ iteration 8840/ 159576 | consumed samples: 506000 | elapsed time per iteration (ms): 23751.6 | learning rate: 6.000E-05 | global batch size: 176 | lm loss: 7.102602E+00 | loss scale: 1024.0 | grad norm: 15258.781 | num zeros: 0.0 | number of skipped iterations: 0 | number of nan iterations: 0 |
66702
+ time (ms)
66703
+ [2021-09-25 19:10:38] PULSE: tr8-104B is scheduled to start in 19:35:46 (at 2021-09-26T14:46:25) (1188168 on 'gpu_p13' partition)
66704
+ [2021-09-25 19:10:38] PULSE: tr8-104B is running for 14:43:37 since 2021-09-25T04:27:01 (1185639_1 on 'gpu_p13' partition (r6i5n[7-8],r6i6n0,r7i2n[4-5],r7i3n2,r7i6n[2-4],r7i7n[7-8],r8i0n[0,2-3,5-8],r8i1n[0,2-4],r8i2n8,r8i3n[0-2],r8i5n[3-4],r8i7n[3-8],r9i0n[0-5],r9i1n[0-3],r9i2n[3-6,8],r9i3n[0-1,7-8],r9i4n[0-3],r9i5n[3-8],r9i6n0)