Comment on Chatbots Make Terrible Doctors, New Study Finds
SuspciousCarrot78@lemmy.world 6 days ago
So, I can speak to this a little bit, as it touches two domains I’m involved it. TL;DR - LLMs bullshit and are unreliable, but there’s a way to use them in this domain as a force multiplier of sorts.
In one; I’ve created a python router that takes my (deidentified) clinical notes, extract and compacts input and creates a summary, then -
-
benchmarks the summary against my (user defined) gold standard and provides management plan (again, based on user defined database).
-
this is then dropped into my on device LLM for light editing and polishing to condense, which I then eyeball, correct and then escalate to supervisor for review.
Additionally, the llm generated note can be approved / denied by the python router, in the first instance based on certain policy criteria I’ve defined.
It can also suggest probable DDX based on my database (which are .CSV based)
Finally, if the llm output fails policy check, the router tells me why it failed and just says “go look at the prior summary and edit it yourself”.
This three step process takes the tedium of paperwork from 15-20 mins to 1 minute generation, 2 mins manual editing.
The reason why this is interesting:
All of this runs within the llm (it calls / invokes the python tooling via >> command) and is 100% deterministic; no llm jazz until the final step, which the router can outright reject and is user auditble anyway.
Ive found that using a fairly “dumb” llm (Qwen2.5-1.5B), with settings dialed down, produces consistently solid final notes (2 out of 3 are graded as passed by router invoking policy document and checking output). Its too dumb to jazz, which is useful in this instance.
Would I trust the LLM, end to end? Well, I’d trust my system, approx 80% of the time. I wouldn’t trust ChatGPT … even though its been more right than wrong in similar tests.
realitista@lemmus.org 6 days ago
Interesting. What technology are you using for this pipeline?
SuspciousCarrot78@lemmy.world 6 days ago
Depends which bit you mean specifically.
The “router” side is a offshoot of a personal project. It’s python scripting and a few other tricks, such as JSON files etc. Full project details for that here
github.com/BobbyLLM/llama-conductor
The tech stack itself:
realitista@lemmus.org 6 days ago
Thanks it’s really interesting to see some real work applications and implementations of AI for practical workloads.
SuspciousCarrot78@lemmy.world 6 days ago
Very welcome :)
As it usually goes with these things, I built it for myself then realised it might have actual broader utility. We shall see!