Philosophy (MA) Psychology (BA)
Pro-AI Vegetarian Misanthrope Cynic Agnostic/Atheist
It is precisely because the users of social media are so so awful that we use it.
Philosophy (MA) Psychology (BA)
Pro-AI Vegetarian Misanthrope Cynic Agnostic/Atheist
It is precisely because the users of social media are so so awful that we use it.
The article states: "ChatGPT-4o performed best with 84.6% validity"
In response to your point: I am mainly interested in probabilistic reliability - if it gives the correct answer 99.9% of the time, it is clearly superior to the vast majority of human beings (with, perhaps, the exception of the best specialists in the most obscure niches) - especially given the sheer breadth of topics is can reliability answer questions on.
Interestingly, my question "What was India like before the British arrived?" produces consistently biased and misleading answers. Though I haven't asked it for the new model.
It is reasonable to assume that the GPT 5.5 on thinking mode has significantly reduced the error rate.
It is also worth noting that the error rate when it comes to diagnosis amongst real doctors is estimated to be around 5%
Admittedly a quite old study: Singh, H., Meyer, A. N. D., & Thomas, E. J. (2014). The frequency of diagnostic errors in outpatient care: Estimations from three large observational studies involving US adult populations. BMJ Quality & Safety, 23(9), 727–731. https://doi.org/10.1136/bmjqs-2013-002627%E2%81%A0%EF%BF%BD