What researchers and scientists say.
RootSource runs on the same retrieval and citation infrastructure behind ExtensionBot, the Extension Foundation's public agricultural chatbot. Here's what independent researchers, a multi-industry AI benchmark, and Extension weed scientists found when they tested ExtensionBot against real questions.
Ranked among the top agriculture-specific models in the first AgriBench results
The Extension Foundation is a founding member of the AgriBench Consortium, housed at the University of Illinois Center for Digital Agriculture alongside Bayer Crop Science, John Deere, Microsoft, and Digital Green, with support from the Gates Foundation and USDA-NIFA. The consortium's first published leaderboard scored ExtensionBot's answers blind against 26 other chatbots, including general-purpose models like ChatGPT and Gemini.
Narrative answer quality across agricultural topics, scored blind by the AgriBench Consortium against 26 other chatbots — nothing about the questions was tailored to ExtensionBot.
This round of AgriBench didn't score citations or hallucination rate — the two things ExtensionBot is built around. That's the strength this benchmark didn't get to credit.
Even without credit for sourcing, it placed ahead of most agriculture-specific tools and competitive with general-purpose models built at far larger scale.
"These results validate the competitiveness of ExtensionBot relative to major commercial and foundation-model systems, particularly in areas where Extension has strong, research-based content." — Extension Foundation, Connect, December 2025
The only chatbot that consistently showed its sources
Four practicing Extension weed scientists tested ExtensionBot against ChatGPT, DeepSeek, and Google Gemini on 24 realistic farmer questions spanning integrated weed management, herbicide selection, resistance, and weed biology.
24 real-world weed management questions — some simple, some multi-step, some with the misspellings a real farmer would type — scored excellent to poor by four Extension weed scientists.
Herbicide resistance and herbicide-specific management — the two categories where a confidently wrong answer costs a grower an entire season.
ChatGPT and DeepSeek rarely cited sources at all. When researchers checked an uncited claim from DeepSeek on herbicide cross-resistance, it was simply wrong.
"Only ExtensionBot consistently linked to the sources of its information, allowing our testers to learn more about a topic, as well as verify the bot's accuracy." — Emily Unglesbee, Grow IWM, March 2026
Asked the same question 1,000 times, ChatGPT gave 1,000 different answers
Researchers from Texas A&M AgriLife Extension and Oklahoma State University tested three OpenAI models on a basic stocking-rate question — asked identically 1,000 times per model, for two different counties. Depending on which GPT version answered, accuracy on the exact same question swung from under 1% to over 99%.
"Is 5 acres per cow/calf pair a sufficient stocking rate?" — asked 1,000 times per model, for a semi-arid Texas county and a wetter Florida county with very different correct answers.
GPT-4 and GPT-4o essentially inverted each other's answers between model versions — both delivered with the same fluent confidence, only one actually consistent with reality.
ExtensionBot answered correctly for the region it had data on, and plainly lacked coverage for the region it didn't — rather than guessing either way.
"Specialized AI platforms trained on region-specific extension publications can offer more accurate and context-relevant advice to producers compared to generalized AI tools like ChatGPT." — Prestegaard-Wilson & Vitale, Animal Frontiers, January 2025
"This is not an effort to replace people at all, but a way to broaden our reach and be able to handle more volume of questions than we could handle otherwise." — Jayson Lusk, Dean of Agriculture, Oklahoma State University, Brownfield Ag News, May 2024
ExtensionBot is backed by a partnership of more than 20 land-grant universities, the Extension Foundation, and the USDA — the same network RootSource's retrieval layer draws from.
Want to see how it holds up against your own use case?
A scoped pilot is the fastest way to know if this fits — measured against your own questions, not a published benchmark.