I wanted a chatbot on the site
A friend asked: what if a chatbot could search the site? That got me thinking.
Justin Chai · 2 October 2026 · 2 min readIt started with me wanting a chatbot on the site. A friend who saw the site said: what if you could use a chatbot to search? So that got me thinking.
I asked Claude how it was done. I tend to host on Cloudflare, so I knew there would definitely be some free Neurons I could use, with OpenAI's open model or Llama. We ran tests and found them somewhat sufficient. (I'm trying to be conservative in what I spend, hah.) But then I realised it wasn't reliable.
Pivoting to Gemini, I found that Gemini had 3.5 Flash-Lite, with a decent free allowance. But it was quite slow, because the model has to take in everything on the site before it can decide which pieces answer the question.
Then I realised there was Jev, a decision model from TypeSafe, which can help with the decision. If it can already decide which articles match the question, it can send just those articles and Gemini can summarise. That is faster than feeding Gemini all of the information. This also scales better as the site grows. The more links I add, the longer Gemini would take to read them all.
And with Jev, it can also decide that a question doesn't match the purpose of Ask, and immediately give a standard response.
Testing
So we had to test Jev. The concept worked, and the results came back faster: the typical wait fell from 7.5 seconds to 1.7.
| Gemini alone | Jev, then Gemini | |
|---|---|---|
| Typical wait per question | 7.5 s | 1.7 s |
| Gemini calls for 20 questions | 40 | 16 |
| Gemini tokens for 20 questions | 138,408 | 23,267 |
| Cost per question at list prices | $0.0024 | $0.0009 |
| Real questions wrongly turned away | 2 | 0 |
| Useful answers (of 16 answerable) | 13 | 15 |
I'm surprised at how much time is saved. On top of that, you are also saving on tokens: Gemini used 83 percent fewer, about 1,200 a question instead of 6,900. That can help immensely where token usage is restricted.
If you are looking at a production-ready environment where you are paying for API calls, Jev is substantially cheaper than a typical API you might use. Jev still reads the index for every question, but at list prices that costs about $0.0003. Having Gemini 3.5 Flash-Lite do the same picking costs about $0.0018. For the whole answer, a question came to about $0.0024 with Gemini alone and about $0.0009 with Jev in front. You do get savings at the end of the day.
My takeaways
Ultimately, I achieved my aim. I always wanted a chatbot, but I didn't really want to learn a proprietary system to set up rules, if-then setups and workflows. At this point, LLMs are smart enough to answer questions. Any questions.
As long as you set up the necessary guardrails.
I would add this. If a language model in your setup is choosing what to read, or deciding whether to answer at all, a decision model in that spot might help to speed things up.
You can try Ask here. I call him... Wevy.