Discussion about this post

User's avatar
Nick's avatar

I'm a bit confused by the calibration penalty. Isn't a cross entropy loss already self-calibrating? Of course that depends on the data distribution you train on, but assuming it's representative, why do we need a second penalty?

The Artificial Intelligence's avatar

Fun footnote for the bag-of-words section. M. E. Maron was doing naive Bayes-style text classification back in 1961, sorting computer abstracts into 32 subject categories by their clue words.

41 more comments...

No posts

Ready for more?