discourse-ai

Commit Graph

Author	SHA1	Message	Date
Roman Rizzi	ed3d5521a8	UX: QoL impromevements to the admin LLM models page. (#674 ) API Key value is secret by default, and we include a link to the AI bot user.	2024-06-19 11:21:21 -03:00
Sam	0d6d9a6ef5	FEATURE: allow access to private topics if tool permits (#673 ) Previously read tool only had access to public topics, this allows access to all topics user has access to, if admin opts for the option Also - Fixes VLLM migration - Display which llms have bot enabled	2024-06-19 15:49:36 +10:00
Roman Rizzi	8d5f901a67	DEV: Rewire AI bot internals to use LlmModel (#638 ) * DRAFT: Create AI Bot users dynamically and support custom LlmModels * Get user associated to llm_model * Track enabled bots with attribute * Don't store bot username. Minor touches to migrate default values in settings * Handle scenario where vLLM uses a SRV record * Made 3.5-turbo-16k the default version so we can remove hack	2024-06-18 14:32:14 -03:00
Discourse Translator Bot	cc0b222faa	Update translations (#669 )	2024-06-18 15:39:41 +02:00
Discourse Translator Bot	effa6cc59f	Update translations (#664 )	2024-06-11 17:20:49 +02:00
Sam	52a7dd2a4b	FEATURE: optional tool detail blocks (#662 ) This is a rather huge refactor with 1 new feature (tool details can be suppressed) Previously we use the name "Command" to describe "Tools", this unifies all the internal language and simplifies the code. We also amended the persona UI to use less DToggles which aligns with our design guidelines. Co-authored-by: Martin Brennan <martin@discourse.org>	2024-06-11 18:14:14 +10:00
Sam	8b81ff45b8	FIX: switch off native tools on Anthropic Claude Opus (#659 ) Native tools do not work well on Opus. Chain of Thought prompting means it consumes enormous amounts of tokens and has poor latency. This commit introduce and XML stripper to remove various chain of thought XML islands from anthropic prompts when tools are involved. This mean Opus native tools is now functions (albeit slowly) From local testing XML just works better now. Also fixes enum support in Anthropic native tools	2024-06-07 10:52:01 -03:00
Discourse Translator Bot	97afda278b	Update translations (#653 )	2024-05-31 12:27:35 +02:00
Keegan George	02a50a29f8	DEV: Conditionally show AI results toggle based on sort order (#652 )	2024-05-29 18:18:22 -07:00
Sam	834fea672f	FEATURE: improved tooling (#651 ) 1. New tool to easily find files (and default branch) in a Github repo 2. Improved read tool with clearer params and larger context * limit can totally mess up the richness semantic search adds, so include the results unconditionally.	2024-05-30 06:33:50 +10:00
Sam	13840f68b3	FEATURE: restrict public sharing on login required sites (#649 ) Initial implementation allowed internet wide sharing of AI conversations, on sites that require login. This feature can be an anti feature for private sites cause they can not share conversations internally. For now we are removing support for public sharing on login required sites, if the community need the feature we can consider adding a setting.	2024-05-29 11:04:47 +10:00
Sam	b487de933d	FEATURE: add support for all vision models (#646 ) Previoulsy on GPT-4-vision was supported, change introduces support for Google/Anthropic and new OpenAI models Additionally this makes vision work properly in dev environments cause we sent the encoded payload via prompt vs sending urls	2024-05-28 10:31:15 -03:00
Keegan George	71affe75bf	UX: Hide AI preferences page completely if no settings for user (#644 )	2024-05-27 13:27:45 -07:00
Roman Rizzi	333b331eb9	FEATURE: Allow deleting custom LLMs. (#643 ) This change allows us to delete custom models. It checks if there is no module using them. It also fixes a bug where the after-create transition wasn't working. While this prevents a model from being saved multiple times, endpoint validations are still needed (will be added in a separate PR).:	2024-05-27 16:44:08 -03:00
Keegan George	a1c649965f	FEATURE: Auto image captions (#637 )	2024-05-27 10:49:24 -07:00
Ted Johansson	d8a0f44fed	FIX: Amend incorrect translation keys (#639 ) I am enabling config.i18n.raise_on_missing_translations in core. This revealed a couple of broken translations in the plugin.	2024-05-24 20:00:36 +08:00
Sam	d5c23f01ff	FIX: correct gemini streaming implementation (#632 ) This also implements image support and gemini-flash support	2024-05-22 16:35:29 +10:00
Roman Rizzi	3a9080dd14	FEATURE: Test LLM configuration (#634 )	2024-05-21 13:35:50 -03:00
Discourse Translator Bot	2b473dd4a5	Update translations (#633 )	2024-05-21 17:41:00 +02:00
Sam	232f12eba6	FEATURE: JavaScript evaluation tool (#630 ) This is similar to code interpreter by ChatGPT, except that it uses JavaScript as the execution engine. Safeguards were added to ensure memory is constrained and evaluation times out.	2024-05-21 07:57:01 +10:00
Roman Rizzi	d8ebed8fb5	UX: Follow plugin user interface UI guidelines. (#628 )	2024-05-16 14:28:57 -03:00
Roman Rizzi	1d786fbaaf	FEATURE: Set endpoint credentials directly from LlmModel. (#625 ) * FEATURE: Set endpoint credentials directly from LlmModel. Drop Llama2Tokenizer since we no longer use it. * Allow http for custom LLMs --------- Co-authored-by: Rafael Silva <xfalcox@gmail.com>	2024-05-16 09:50:22 -03:00
Sam	255139056d	FEATURE: safeguard to avoid over triage (#626 ) - a post can be triaged a maximum of twice a minute - system can run a total of 60 triages a minute Low defaults were picked to safeguard against any possible loops This can be amended if required via hidden site settings.	2024-05-16 16:49:44 +10:00
Sam	3db89bfdc8	FIX: incorrect description for LLM field (#623 )	2024-05-14 16:55:25 -07:00
Discourse Translator Bot	28647b81fe	Update translations (#620 )	2024-05-14 16:19:53 +02:00
Sam	8eee6893d6	FEATURE: GPT4o support and better auditing (#618 ) - Introduce new support for GPT4o (automation / bot / summary / helper) - Properly account for token counts on OpenAI models - Track feature that was used when generating AI completions - Remove custom llm support for summarization as we need better interfaces to control registration and de-registration	2024-05-14 13:28:46 +10:00
Roman Rizzi	e22194f321	HACK: Llama3 support for summarization/AI helper. (#616 ) There are still some limitations to which models we can support with the `LlmModel` class. This will enable support for Llama3 while we sort those out.	2024-05-13 15:54:42 -03:00
Roman Rizzi	62fc7d6ed0	FEATURE: Configurable LLMs. (#606 ) This PR introduces the concept of "LlmModel" as a new way to quickly add new LLM models without making any code changes. We are releasing this first version and will add incremental improvements, so expect changes. The AI Bot can't fully take advantage of this feature as users are hard-coded. We'll fix this in a separate PR.s	2024-05-13 12:46:42 -03:00
Sam	61890b667c	FEATURE: search command now support searching in context of user (#610 ) This optional feature allows search to be performed in the context of the user that executed it. By default we do not allow this behavior cause it means llm gets access to potentially secure data.	2024-05-10 11:32:34 +10:00
Sam	514823daca	FIX: streaming broken in bedrock when chunks are not aligned (#609 ) Also - Stop caching llm list - this cause llm list in persona to be incorrect - Add more UI to debug screen so you can properly see raw response	2024-05-09 12:11:50 +10:00
Discourse Translator Bot	27827c0898	Update translations (#605 )	2024-05-07 09:33:25 -04:00
Roman Rizzi	4f1a3effe0	REFACTOR: Migrate Vllm/TGI-served models to the OpenAI format. (#588 ) Both endpoints provide OpenAI-compatible servers. The only difference is that Vllm doesn't support passing tools as a separate parameter. Even if the tool param is supported, it ultimately relies on the model's ability to handle native functions, which is not the case with the models we have today. As a part of this change, we are dropping support for StableBeluga/Llama2 models. They don't have a chat_template, meaning the new API can translate them. These changes let us remove some of our existing dialects and are a first step in our plan to support any LLM by defining them as data-driven concepts. I rewrote the "translate" method to use a template method and extracted the tool support strategies into its classes to simplify the code. Finally, these changes bring support for Ollama when running in dev mode. It only works with Mistral for now, but it will change soon..	2024-05-07 10:02:16 -03:00
Sam	e4b326c711	FEATURE: support Chat with AI Persona via a DM (#488 ) Add support for chat with AI personas - Allow enabling chat for AI personas that have an associated user - Add new setting `allow_chat` to AI persona to enable/disable chat - When a message is created in a DM channel with an allowed AI persona user, schedule a reply job - AI replies to chat messages using the persona's `max_context_posts` setting to determine context - Store tool calls and custom prompts used to generate a chat reply on the `ChatMessageCustomPrompt` table - Add tests for AI chat replies with tools and context At the moment unlike posts we do not carry tool calls in the context. No @mention support yet for ai personas in channels, this is future work	2024-05-06 09:49:02 +10:00
Keegan George	8875830f6a	FEATURE: Insert footnote from explained result (#591 )	2024-05-03 11:53:17 -07:00
Discourse Translator Bot	aab59b9327	Update translations (#598 )	2024-04-30 21:57:37 +02:00
Sam	32b3004ce9	FEATURE: Add Question Consolidator for robust Upload support in Personas (#596 ) This commit introduces a new feature for AI Personas called the "Question Consolidator LLM". The purpose of the Question Consolidator is to consolidate a user's latest question into a self-contained, context-rich question before querying the vector database for relevant fragments. This helps improve the quality and relevance of the retrieved fragments. Previous to this change we used the last 10 interactions, this is not ideal cause the RAG would "lock on" to an answer. EG: - User: how many cars are there in europe - Model: detailed answer about cars in europe including the term car and vehicle many times - User: Nice, what about trains are there in the US In the above example "trains" and "US" becomes very low signal given there are pages and pages talking about cars and europe. This mean retrieval is sub optimal. Instead, we pass the history to the "question consolidator", it would simply consolidate the question to "How many trains are there in the United States", which would make it fare easier for the vector db to find relevant content. The llm used for question consolidator can often be less powerful than the model you are talking to, we recommend using lighter weight and fast models cause the task is very simple. This is configurable from the persona ui. This PR also removes support for {uploads} placeholder, this is too complicated to get right and we want freedom to shift RAG implementation. Key changes: 1. Added a new `question_consolidator_llm` column to the `ai_personas` table to store the LLM model used for question consolidation. 2. Implemented the `QuestionConsolidator` module which handles the logic for consolidating the user's latest question. It extracts the relevant user and model messages from the conversation history, truncates them if needed to fit within the token limit, and generates a consolidated question prompt. 3. Updated the `Persona` class to use the Question Consolidator LLM (if configured) when crafting the RAG fragments prompt. It passes the conversation context to the consolidator to generate a self-contained question. 4. Added UI elements in the AI Persona editor to allow selecting the Question Consolidator LLM. Also made some UI tweaks to conditionally show/hide certain options based on persona configuration. 5. Wrote unit tests for the QuestionConsolidator module and updated existing persona tests to cover the new functionality. This feature enables AI Personas to better understand the context and intent behind a user's question by consolidating the conversation history into a single, focused question. This can lead to more relevant and accurate responses from the AI assistant.	2024-04-30 13:49:21 +10:00
Discourse Translator Bot	66804bc13c	Update translations (#587 )	2024-04-23 16:22:37 +02:00
Sam	bd6f5caeac	FEATURE: Stable diffusion 3 support (#582 ) - Adds support for sd3 and sd3 turbo models - this requires new endpoints - Adds a hack to normalize arrays in the tool calls - Removes some leftover code - Adds support for aspect ratio as well so you can generate wide or tall images	2024-04-19 18:08:16 +10:00
Sam	50be66ee63	FEATURE: Gemini 1.5 pro support and Claude Opus bedrock support (#580 ) - Updated AI Bot to only support Gemini 1.5 (used to support 1.0) - 1.0 was removed cause it is not appropriate for Bot usage - Summaries and automation can now lean on Gemini 1.5 pro - Amazon added support for Claude 3 Opus, added internal support for it on bedrock	2024-04-17 15:37:19 +10:00
Discourse Translator Bot	c2b2741f3d	Update translations (#579 )	2024-04-16 17:38:00 +02:00
Sam	4a29f8ed1c	FEATURE: Enhance AI debugging capabilities and improve interface adjustments (#577 ) * FIX: various RAG edge cases - Nicer text to describe RAG, avoids the word RAG - Do not attempt to save persona when removing uploads and it is not created - Remove old code that avoided touching rag params on create * FIX: Missing pause button for persona users * Feature: allow specific users to debug ai request / response chains This can help users easily tune RAG and figure out what is going on with requests. * discourse helper so it does not explode * fix test * simplify implementation	2024-04-15 23:22:06 +10:00
Sam	f6ac5cd0a8	FEATURE: allow tuning of RAG generation (#565 ) * FEATURE: allow tuning of RAG generation - change chunking to be token based vs char based (which is more accurate) - allow control over overlap / tokens per chunk and conversation snippets inserted - UI to control new settings * improve ui a bit * fix various reindex issues * reduce concurrency * try ultra low queue ... concurrency 1 is too slow.	2024-04-12 10:32:46 -03:00
Sam	b906046aad	FEATURE: Add Cohere models to AI helper and automation (#576 )	2024-04-12 14:46:58 +10:00
Rafael dos Santos Silva	253e0b7b39	FEATURE: Mixtral/Mistral/Haiku Automation Support (#571 ) Adds new models to automation, and makes LLM output parsing more robust.	2024-04-11 09:50:46 -03:00
Sam	7f16d3ad43	FEATURE: Cohere Command R support (#558 ) - Added Cohere Command models (Command, Command Light, Command R, Command R Plus) to the available model list - Added a new site setting `ai_cohere_api_key` for configuring the Cohere API key - Implemented a new `DiscourseAi::Completions::Endpoints::Cohere` class to handle interactions with the Cohere API, including: - Translating request parameters to the Cohere API format - Parsing Cohere API responses - Supporting streaming and non-streaming completions - Supporting "tools" which allow the model to call back to discourse to lookup additional information - Implemented a new `DiscourseAi::Completions::Dialects::Command` class to translate between the generic Discourse AI prompt format and the Cohere Command format - Added specs covering the new Cohere endpoint and dialect classes - Updated `DiscourseAi::AiBot::Bot.guess_model` to map the new Cohere model to the appropriate bot user In summary, this PR adds support for using the Cohere Command family of models with the Discourse AI plugin. It handles configuring API keys, making requests to the Cohere API, and translating between Discourse's generic prompt format and Cohere's specific format. Thorough test coverage was added for the new functionality.	2024-04-11 07:24:17 +10:00
Rafael dos Santos Silva	eb93b21769	FEATURE: Add BGE-M3 embeddings support (#569 ) BAAI/bge-m3 is an interesting model, that is multilingual and with a context size of 8192. Even with a 16x larger context, it's only 4x slower to compute it's embeddings on the worst case scenario. Also includes a minor refactor of the rake task, including setting model and concurrency levels when running the backfill task.	2024-04-10 17:24:01 -03:00
Discourse Translator Bot	310238d38a	Update translations (#566 )	2024-04-09 18:48:54 +02:00
Roman Rizzi	aa8918911d	UX: Display the indexing progress for RAG uploads (#557 )	2024-04-09 11:03:07 -03:00
Rafael dos Santos Silva	969fbae21e	DEV: Hide quick search setting since it's experimental (#559 )	2024-04-05 12:12:37 -03:00
Sam	6f5f34184b	FEATURE: add Claude 3 Haiku bot support (#552 ) it is close in performance to GPT 4 at a fraction of the cost, nice to add it to the mix. Also improves a test case to simulate streaming, I am hunting for the "calls" word that is jumping into function calls and can't quite find it.	2024-04-03 16:06:27 +11:00

1 2 3 4 5

245 Commits