AI Megathread
-
This English professor asserting that em-dashes are the biggest tell-tale sign of her students using AI is what I’m talking about when I complain about people picking on the dashes.
-
bring back downvotes
-
@Rathenhope said in AI Megathread:
two potential clients (the biggest we’d have) both went “we consider any use of AI to be high risk and we don’t want client information anywhere near it”
Just casually send me the names of these clients so I can get our CEO to try and sell to them, I really want to see that metaphorical ice bucket tipped all over him when he hears that.
-
FYI, AI detectors are still kinda crappy.
Can AI Detectors Be Trusted? The Authors Guild Put Five of Them to the Test
“Our test confirms that while a couple of AI detection tools accurately identify human-authored text, some commonly used consumer-facing AI detection tools are wildly inaccurate, which presents a major risk for authors. Moreover, these tools change constantly—updated models, shifting benchmarks, evolving AI outputs—and their accuracy at any given moment cannot be assumed.”
Pangram Flagged My Own Writing as AI
Regarding false positive benchmarks: “that benchmark tests pure human text and pure AI text under controlled conditions. It doesn’t describe real-world use cases”
-
I don’t agree with your conclusions. From the same article:
The results varied widely, and in some cases, dramatically.
Pangram and Originality.ai were the most reliable performers. Pangram returned 0 percent across all ten articles. Originality.ai returned 0 percent on eight of ten, with 1 percent on the remaining two. Both tools correctly identified every piece as human written.
Grammarly performed nearly as well, returning 0 percent on eight articles and flagging two at 7 percent and 9 percent respectively—low enough that neither would likely trigger concern in practice.3 out of 5 did okay? They are tools to use alongside human reasoning. But demonstrably they aren’t garbage.
-
@Tez said in AI Megathread:
3 out of 5 did okay? They are tools to use alongside human reasoning. But demonstrably they aren’t garbage.
One of the tools you’re saying did “okay” in the first article is the same one the second article raises concerns about. Do they sometimes work? Sure. Are they reliable across a wide variety of use cases? No.
-
@Faraday IDK, I’d just continue to call them imperfect tools. That just puts it at 1 false flag in 11 samples instead of it’s 0 false flags out of 10 samples. Still not garbage.
I think there are real issues in these tools, but I think it is a mistake to write it all off as garbage. Use it as a tool alongside human judgment. Try to use the best tool you can.
-
@Tez said in AI Megathread:
Use it as a tool alongside human judgment.
Obviously this is the correct course, but basically every institutional use that I have personally seen has been lacking in the latter section.
Which is a problem in terms of work culture, not the tool, yes but when you’re relying on the tool to make decisions that directly impact a person’s likelihood of getting a degree, or getting/keeping a job? I don’t think we should be using, much less relying on, imperfect tools for in those instances.
-
@Pavel said in AI Megathread:
Obviously this is the correct course, but basically every institutional use that I have personally seen has been lacking in the latter section.
Which is a problem in terms of work culture, not the tool, yes but when you’re relying on the tool to make decisions that directly impact a person’s likelihood of getting a degree, or getting/keeping a job? I don’t think we should be using, much less relying on, imperfect tools for in those instances.
This is my problem, yes. However, I attribute it to the tool as much as the work culture.
These things portray themselves as far more reliable than they actually are in practice. The UX also lends toward this reliance, displaying results like “31% likely to be AI”. That, again, is giving the misleading impression of a highly precise and reliable tool. And unlike a plagiarism detector, which can point you to the original, these tools are much more opaque.
You don’t have to look far on substack to find tons of authors (pissed off about the recent addition of Pangram to the site) pointing out false positives in their own work. Even a comparatively low false-positive rate can create massive implications when used at large scales.
-
@Aria said in AI Megathread:
I’m pretty sure we’re all just living in the plot of Wall-E now and I hate it here.
Wall-E and Idiocracy. Yup.