Personally I've never seen a training protocol for BERT models [1] that I trust enough that I could build it into an automated system that updates itself. On the other hand if i want to do classification or something like that it is straightforward to pool and train a classical classifier or, for harder problems, train an LSTM on the token-by-token output.
If you are trying to change the way a seq-to-seq works that's a different matter.
[1] people use phrases like "catastrophic forgetting" but my take on it is "if your model trainer doesn't support early stopping I don't want to start"
To do what?
Personally I've never seen a training protocol for BERT models [1] that I trust enough that I could build it into an automated system that updates itself. On the other hand if i want to do classification or something like that it is straightforward to pool and train a classical classifier or, for harder problems, train an LSTM on the token-by-token output.
If you are trying to change the way a seq-to-seq works that's a different matter.
[1] people use phrases like "catastrophic forgetting" but my take on it is "if your model trainer doesn't support early stopping I don't want to start"