This is a personal project to explore ethical applications of LLMs, it is privacy focused, utilising a local LLM, lightweight, non tracking and providing a public good.
The results here are promising, however, the quality of answer is not in line with large expensive models. Recall and refusal are quite strong, Blaai is far more likely than not to share an authoritative source where available and admit it's limitations where not. However, the generated answers themselves are frequently incomplete and sometimes a little strange.
| Metric | Mean | yes | partial | no |
|---|---|---|---|---|
| groundedness | 0.53 | 2 | 15 | 1 |
| correctness | 0.42 | 1 | 13 | 4 |