A Consumer Perspective on Evaluation
What happens if we think of evaluations as a way of helping users choose the best NLP tech for their needs?
What happens if we think of evaluations as a way of helping users choose the best NLP tech for their needs?
A summary of students who have gotten PhDs under my supervision.
A visitor recently asked me what NLG research topics I found most interesting and exciting. Great question, I’ve written here an expanded version of what I told him.
My response to Goldberg’s “adversarial review” of some research on using deep learning in NLG
I’m planning to do a systematic review of the validity of BLEU, and am very keen to get comments and suggestions on study design from others!
A short travelogue about a holiday cycling trip I did in May 2017 (nothing to do with NLG!)
I’m looking for a PhD student to work on Advanced Data Storytelling!
Some explanation and advice about regression to mean, which is a statistical phenomena that can impact NLG evaluations.
People who use corpora to build NLG systems need to understand what is in the corpora. The widely used Weathergov corpus, for example, probably contains computer-generated texts rather than human-written texts. So learning from it is essentially reverse-engineering a rule-based NLG system.
I am really dubious about evaluations based on BLEU and other metrics. I explain why, and also give advice on best practice for people who are committed to using metrics