Comment

Comment on Chatbots Make Terrible Doctors, New Study Finds

SuspciousCarrot78@lemmy.world ⁨3⁩ ⁨months⁩ ago

So, I can speak to this a little bit, as it touches two domains I’m involved it. TL;DR - LLMs bullshit and are unreliable, but there’s a way to use them in this domain as a force multiplier of sorts.

In one; I’ve created a python router that takes my (deidentified) clinical notes, extract and compacts input and creates a summary, then -

benchmarks the summary against my (user defined) gold standard and provides management plan (again, based on user defined database).
this is then dropped into my on device LLM for light editing and polishing to condense, which I then eyeball, correct and then escalate to supervisor for review.

Additionally, the llm generated note can be approved / denied by the python router, in the first instance based on certain policy criteria I’ve defined.

It can also suggest probable DDX based on my database (which are .CSV based)

Finally, if the llm output fails policy check, the router tells me why it failed and just says “go look at the prior summary and edit it yourself”.

This three step process takes the tedium of paperwork from 15-20 mins to 1 minute generation, 2 mins manual editing.

The reason why this is interesting:

All of this runs within the llm (it calls / invokes the python tooling via >> command) and is 100% deterministic; no llm jazz until the final step, which the router can outright reject and is user auditble anyway.

Ive found that using a fairly “dumb” llm (Qwen2.5-1.5B), with settings dialed down, produces consistently solid final notes (2 out of 3 are graded as passed by router invoking policy document and checking output). Its too dumb to jazz, which is useful in this instance.

Would I trust the LLM, end to end? Well, I’d trust my system, approx 80% of the time. I wouldn’t trust ChatGPT … even though its been more right than wrong in similar tests.

source

Sort:hotnew top

realitista@lemmus.org ⁨3⁩ ⁨months⁩ ago
Interesting. What technology are you using for this pipeline?

source
- SuspciousCarrot78@lemmy.world ⁨3⁩ ⁨months⁩ ago
  Depends which bit you mean specifically.
  
  The “router” side is a offshoot of a personal project. It’s python scripting and a few other tricks, such as JSON files etc. Full project details for that here
  
  github.com/BobbyLLM/llama-conductor
  
  The tech stack itself:
  
  llama.cpp
  
  Qwen 2.5-1.5 GGUF base (by memory, 5 bit quant from HF Alibaba repository)
  
  The python router (more sophisticated version of above)
  
  Policy documents
  
  Front end (OWUI - may migrate to something simpler / more robust)
  
  source
  - realitista@lemmus.org ⁨3⁩ ⁨months⁩ ago
    Thanks it’s really interesting to see some real work applications and implementations of AI for practical workloads.
    
    source
    SuspciousCarrot78@lemmy.world ⁨3⁩ ⁨months⁩ ago
    Very welcome :)
    
    As it usually goes with these things, I built it for myself then realised it might have actual broader utility. We shall see!
    
    source