That's a very interesting paper (and very accessible to anyone with a stats/data mining background). I went back and read Jason Baldridge's intro, which is excellent
It seems you didn't attempt to fingerprint for misspellings, among the variables on pdf p 5. Also, curious why did you need to up the dataset to exactly 100k with the 5.7k.
http://ata-s12.utcompling.com/schedule/ATA-Authorship%20Attr...
It seems you didn't attempt to fingerprint for misspellings, among the variables on pdf p 5. Also, curious why did you need to up the dataset to exactly 100k with the 5.7k.