There is a strange and seemingly impossible to define boundary between when LLMs like ChatGPT are useful and when they are useless. Where they are useless is easy to find: one example being when I asked Microsoft Copilot what the precise meaning of the “1 of 30 responses” that it gives at the bottom of each response is. One would expect that it would be very good at answering questions about itself, but I was unable to get a sensible answer out of it. My best guess is that it is a memory limit on the session.
But then it surprises you with responses that are really quite extraordinary. For example, I found the following unintelligible equations in a paper I was reading:

When I asked ChatGPT what equation (2.92) means it replied that it contains non-standard nomenclature and then gave its interpretation of it. What is really extraordinary is that, in the context of the paper the equation came from, I think that ChatGPT’s response is probably correct! And I used the word ‘probably’ advisedly here, because all of the responses of an LLM are just probabilistic random output. What is difficult with LLMs is trying to discern when the random output will be useful and when it will be useless.
The chat I had with ChatGPT is below:

Postscript
It occurred to me that the reason that MS Copilot was unable to give a sensible answer is that my question was using the second person to refer to Copilot, which would not correspond to any of the data that it was trained on. I therefore tried again with a question that refers to Copilot in the third person and it was then able to give me a good response – See below. It seems that another requirement of generating a good prompt is to not treat it like a person but to always keep in mind the data it is trained on. In this case it would be questions and answers which refer to Copilot in the third person.
Or perhaps it was something else about the prompt that gave the bad output, which then raises the additional question of how much time is it worth spending on learning how to create good prompts? More importantly though, when is a bad response due to the LLM being bad at the task and when is it due to the prompt? One big difference between an LLM and a person is that a person will ask questions to clarify your question, but an LLM will only give clarifying questions if you ask it to. In future I will try the prompt “Ask me a question that will improve your response.” after bad answers to see if that helps.
