discourse-ai

Commit Graph

Author	SHA1	Message	Date
Rafael dos Santos Silva	0d3e6b2726	FIX: Fix ordering of random post embeddings backfill (#965 ) * FIX: Fix ordering of random post embeddings backfill * fix annotations --------- Co-authored-by: Roman Rizzi <rizziromanalejandro@gmail.com>	2024-11-27 17:01:54 -03:00
Roman Rizzi	ef07fcb308	FIX: Skip records without content to classify (#960 )	2024-11-26 15:54:20 -03:00
Roman Rizzi	ddf2bf7034	DEV: Backfill embeddings concurrently. (#941 ) We are adding a new method for generating and storing embeddings in bulk, which relies on `Concurrent::Promises::Future`. Generating an embedding consists of three steps: Prepare text HTTP call to retrieve the vector Save to DB. Each one is independently executed on whatever thread the pool gives us. We are bringing a custom thread pool instead of the global executor since we want control over how many threads we spawn to limit concurrency. We also avoid firing thousands of HTTP requests when working with large batches.	2024-11-26 14:12:32 -03:00
Rafael dos Santos Silva	23193ee6f2	FEATURE: Calculate gists from non hot topics too (#958 ) Also renames some settings to remove 'hot' references.	2024-11-26 13:44:12 -03:00
Roman Rizzi	fbc74c7467	FEATURE: Extend summary backfill to also generate gists (#896 ) Updates default batch size to 0 and max to 10000	2024-11-07 13:40:18 -03:00
Roman Rizzi	9505a8976c	FEATURE: Automatically backfill regular summaries. (#892 ) This change introduces a job to summarize topics and cache the results automatically. We provide a setting to control how many topics we'll backfill per hour and what the topic's minimum word count is to qualify. We'll prioritize topics without summary over outdated ones.	2024-11-04 17:48:11 -03:00
Roman Rizzi	a2b1ea3c63	FEATURE: Fast-track gist regeneration when a hot topic gets a new post (#860 ) * FEATURE: Fast-track gist regeneration when a hot topic gets a new post * DEV: Introduce an upsert-like summarize * FIX: Only enqueue fast-track gist for hot hot hot topics --------- Co-authored-by: Rafael Silva <xfalcox@gmail.com>	2024-10-25 12:38:49 -03:00
Roman Rizzi	6d504ab80d	FEATURE: Make hot topic gists opt-in. (#846 ) This change restricts gists to members of specific groups. It also fixes a bug where other lists could display the gist if available.	2024-10-21 15:15:25 -03:00
Roman Rizzi	e768fa877e	FIX: Don't regenerate up to date gists (#843 )	2024-10-18 18:49:01 -03:00
Roman Rizzi	27b5542357	FEATURE: Generate topic gists for the hot topics list. (#837 ) * Display gists in the hot topics list * Adjust hot topics gist strategy and add a job to generate gists * Replace setting with a configurable batch size * Avoid loading summaries for other topic lists * Tweak gist prompt to focus on latest posts in the context of the OP * Remove serializer hack and rely on core change from discourse/discourse#29291 * Update lib/summarization/strategies/hot_topic_gists.rb Co-authored-by: Rafael dos Santos Silva <xfalcox@gmail.com> --------- Co-authored-by: Rafael dos Santos Silva <xfalcox@gmail.com>	2024-10-18 18:01:39 -03:00
Rafael dos Santos Silva	792703c942	FEATURE: Discord Bot integration (#831 ) This adds support for the a Discord bot that can search in a Discourse instance when invoked via slash commands in Discord Guild channel.	2024-10-16 12:41:18 -03:00
Roman Rizzi	c7acb4a6a0	REFACTOR: Support of different summarization targets/prompts. (#835 ) * DEV: Add summary types * Refactor for different summary types * Use enum for summary types * Update lib/summarization/strategies/topic_summary.rb Co-authored-by: Penar Musaraj <pmusaraj@gmail.com> * Update lib/summarization/strategies/topic_gist.rb Co-authored-by: Penar Musaraj <pmusaraj@gmail.com> * Update lib/summarization/strategies/chat_messages.rb Co-authored-by: Penar Musaraj <pmusaraj@gmail.com> * Fix chat_messages single prompt * Small tweak to the chat summarization prompt --------- Co-authored-by: Penar Musaraj <pmusaraj@gmail.com>	2024-10-15 13:53:26 -03:00
Rafael dos Santos Silva	791fad1e6a	FEATURE: Index embeddings using bit vectors (#824 ) On very large sites, the rare cache misses for Related Topics can take around 200ms, which affects our p99 metric on the topic page. In order to mitigate this impact, we now have several tools at our disposal. First, one is to migrate the index embedding type from halfvec to bit and change the related topic query to leverage the new bit index by changing the search algorithm from inner product to Hamming distance. This will reduce our index sizes by 90%, severely reducing the impact of embeddings on our storage. By making the related query a bit smarter, we can have zero impact on recall by using the index to over-capture N2 results, then re-ordering those N2 using the full halfvec vectors and taking the top N. The expected impact is to go from 200ms to <20ms for cache misses and from a 2.5GB index to a 250MB index on a large site. Another tool is migrating our index type from IVFFLAT to HNSW, which can increase the cache misses performance even further, eventually putting us in the under 5ms territory. Co-authored-by: Roman Rizzi <roman@discourse.org>	2024-10-14 13:26:03 -03:00
Sam	5cbc9190eb	FEATURE: RAG search within tools (#802 ) This allows custom tools access to uploads and sophisticated searches using embedding. It introduces: - A shared front end for listing and uploading files (shared with personas) - Backend implementation of index.search function within a custom tool. Custom tools now may search through uploaded files function invoke(params) { return index.search(params.query) } This means that RAG implementers now may preload tools with knowledge and have high fidelity over the search. The search function support specifying max results specifying a subset of files to search (from uploads) Also - Improved documentation for tools (when creating a tool a preamble explains all the functionality) - uploads were a bit finicky, fixed an edge case where the UI would not show them as updated	2024-09-30 17:27:50 +10:00
Sam	03eccbe392	FEATURE: Make tool support polymorphic (#798 ) Polymorphic RAG means that we will be able to access RAG fragments both from AiPersona and AiCustomTool In turn this gives us support for richer RAG implementations.	2024-09-16 08:17:17 +10:00
Sam	a48acc894a	FEATURE: more accurate and faster titles (#791 ) Previously we waited 1 minute before automatically titling PMs The new change introduces adding a title immediately after the the llm replies Prompt was also modified to include the LLM reply in title suggestion. This helps situation like: user: tell me a joke llm: a very funy joke about horses Then the title would be "A Funny Horse Joke" Specs already covered some auto title logic, amended to also catch the new message bus message we have been sending.	2024-09-03 15:52:20 +10:00
Sam	584753cf60	FIX: we were never reindexing old content (#786 ) * FIX: we were never reindexing old content Embedding backfill contains logic for searching for old content change and then backfilling. Unfortunately it was excluding all topics that had embedding unconditionally, leading to no backfill ever happening. This change adds a test and ensures we backfill. * over select results, this ensures we will be more likely to find ai results when filtered	2024-08-30 14:37:55 +10:00
Keegan George	fdadfa029e	FEATURE: smooth streaming animation for summarization (#778 )	2024-08-29 15:07:07 -07:00
Keegan George	94f6c632bf	DEV: Publish AI Bot PM title update to message bus channel (#781 )	2024-08-29 14:48:44 -07:00
Sam	14443bf890	FIX: more robust summary implementation (#750 ) When navigating between topic we were not correctly resetting internal state for summarization. This leads to a situation where incorrect summaries can be displayed to users and wrong summaries can be displayed. Additionally our controller for grabbing summaries was always streaming results via message bus, which could be delayed when sidekiq is overloaded. We now will return the cached summary right away if it is available direct from REST endpoint.	2024-08-13 08:47:47 -03:00
Keegan George	1d6a6c9f8f	FEATURE: Stream other post helper options (#745 )	2024-08-08 11:32:39 -07:00
Sam	1320eed9b2	FEATURE: move summary to use llm_model (#699 ) This allows summary to use the new LLM models and migrates of API key based model selection Claude 3.5 etc... all work now. --------- Co-authored-by: Roman Rizzi <rizziromanalejandro@gmail.com>	2024-07-04 10:48:18 +10:00
Keegan George	1b0ba9197c	DEV: Add summarization logic from core (#658 )	2024-07-02 08:51:59 -07:00
Loïc Guitaut	dd4e305ff7	DEV: Update rubocop-discourse to version 3.8.0 (#641 )	2024-05-28 11:15:42 +02:00
Sam	d4116ecfac	FEATURE: Add support for contextualizing a DM to a bot (#627 ) This brings the context of the current topic on screen into chat	2024-05-21 17:17:02 +10:00
Sam	e4b326c711	FEATURE: support Chat with AI Persona via a DM (#488 ) Add support for chat with AI personas - Allow enabling chat for AI personas that have an associated user - Add new setting `allow_chat` to AI persona to enable/disable chat - When a message is created in a DM channel with an allowed AI persona user, schedule a reply job - AI replies to chat messages using the persona's `max_context_posts` setting to determine context - Store tool calls and custom prompts used to generate a chat reply on the `ChatMessageCustomPrompt` table - Add tests for AI chat replies with tools and context At the moment unlike posts we do not carry tool calls in the context. No @mention support yet for ai personas in channels, this is future work	2024-05-06 09:49:02 +10:00
Roman Rizzi	283445cf81	FIX: RAG uploader must support multi-file indexing. (#592 ) Updating the editing model's rag_uploads in the editor component broke multi-file uploading. Instead, we'll keep the uploads in the uploader and update the model when we finish. This PR also fast-tracks the initial update so we can show feedback to the user quickly, and allows uploading MD files. Bug reported on https://meta.discourse.org/t/discourse-ai-persona-upload-support/304049/11	2024-04-25 10:48:55 -03:00
Sam	a5e4ab2825	FIX: blank metadata leading to errors (#578 ) blank metadata block in RAG was leading to an error, this handles the edge case	2024-04-17 13:46:40 +10:00
Sam	f6ac5cd0a8	FEATURE: allow tuning of RAG generation (#565 ) * FEATURE: allow tuning of RAG generation - change chunking to be token based vs char based (which is more accurate) - allow control over overlap / tokens per chunk and conversation snippets inserted - UI to control new settings * improve ui a bit * fix various reindex issues * reduce concurrency * try ultra low queue ... concurrency 1 is too slow.	2024-04-12 10:32:46 -03:00
Martin Brennan	bab5e52e38	FIX: Secure/unsecure uploads when sharing AI conversations (#554 ) This commit uses a new plugin modifier introduced in https://github.com/discourse/discourse/pull/26508 to mark all uploads as _not_ secure in shared PM AI conversations. This is so images created by the AI bot (or uploaded by the user) do not end up as broken URLs because of the security requirements around them. This relies on the UpdateTopicUploadSecurity job in core as well, which is fired when an AI conversation is shared or deleted.	2024-04-11 10:00:41 +10:00
Roman Rizzi	aa8918911d	UX: Display the indexing progress for RAG uploads (#557 )	2024-04-09 11:03:07 -03:00
Sam	830cc26075	FEATURE: Add metadata support for RAG (#553 ) * FEATURE: Add metadata support for RAG You may include non indexed metadata in the RAG document by using [[metadata ....]] This information is attached to all the text below and provided to the retriever. This allows for RAG to operate within a rich amount of contexts without getting lost Also: - re-implemented chunking algorithm so it streams - moved indexing to background low priority queue * Baran gem no longer required. * tokenizers is on 4.4 ... upgrade it ...	2024-04-04 11:02:16 -03:00
Roman Rizzi	1f1c94e5c6	FEATURE: AI Bot RAG support. (#537 ) This PR lets you associate uploads to an AI persona, which we'll split and generate embeddings from. When building the system prompt to get a bot reply, we'll do a similarity search followed by a re-ranking (if available). This will let us find the most relevant fragments from the body of knowledge you associated with the persona, resulting in better, more informed responses. For now, we'll only allow plain-text files, but this will change in the future. Commits: * FEATURE: RAG embeddings for the AI Bot This first commit introduces a UI where admins can upload text files, which we'll store, split into fragments, and generate embeddings of. In a next commit, we'll use those to give the bot additional information during conversations. * Basic asymmetric similarity search to provide guidance in system prompt * Fix tests and lint * Apply reranker to fragments * Uploads filter, css adjustments and file validations * Add placeholder for rag fragments * Update annotations	2024-04-01 13:43:34 -03:00
Sam	a03bc6ddec	FEATURE: Share conversations with AI via a URL (#521 ) This allows users to share a static page of an AI conversation with the rest of the world. By default this feature is disabled, it is enabled by turning on ai_bot_allow_public_sharing via site settings Precautions are taken when sharing 1. We make a carbonite copy 2. We minimize work generating page 3. We limit to 100 interactions 4. Many security checks - including disallowing if there is a mix of users in the PM. * Bonus commit, large PRs like this PR did not work with github tool large objects would destroy context Co-authored-by: Martin Brennan <martin@discourse.org>	2024-03-12 16:51:41 +11:00
Sam	484fd1435b	DEV: improve internal design of ai persona and bug fix (#495 ) * DEV: improve internal design of ai persona and bug fix - Fixes bug where OpenAI could not describe images - Fixes bug where mentionable personas could not be mentioned unless overarching bot was enabled - Improves internal design of playground and bot to allow better for non "bot" users - Allow PMs directly to persona users (previously bot user would also have to be in PM) - Simplify internal code Co-authored-by: Martin Brennan <martin@discourse.org>	2024-02-28 16:46:32 +11:00
Sam	3a8d95f6b2	FEATURE: mentionable personas and random picker tool, context limits (#466 ) 1. Personas are now optionally mentionable, meaning that you can mention them either from public topics or PMs - Mentioning from PMs helps "switch" persona mid conversation, meaning if you want to look up sites setting you can invoke the site setting bot, or if you want to generate an image you can invoke dall e - Mentioning outside of PMs allows you to inject a bot reply in a topic trivially - We also add the support for max_context_posts this allow you to limit the amount of context you feed in, which can help control costs 2. Add support for a "random picker" tool that can be used to pick random numbers 3. Clean up routing ai_personas -> ai-personas 4. Add Max Context Posts so users can control how much history a persona can consume (this is important for mentionable personas) Co-authored-by: Martin Brennan <martin@discourse.org>	2024-02-15 16:37:59 +11:00
Rafael dos Santos Silva	fd6fcfdb61	DEV: Increase embeddings backfill job frequency (#453 ) The idea is to increase the frequency so we can run with smaller batch sizes. Big batches cause problems when running backups, so it's better to have shorter but more frequent jobs.	2024-01-31 15:09:39 -03:00
Sam	dcafc8032f	FIX: improve embedding generation (#452 ) 1. on failure we were queuing a job to generate embeddings, it had the wrong params. This is both fixed and covered in a test. 2. backfill embedding in the order of bumped_at, so newest content is embedded first, cover with a test 3. add a safeguard for hidden site setting that only allows batches of 50k in an embedding job run Previously old embeddings were updated in a random order, this changes it so we update in a consistent order	2024-01-31 10:38:47 -03:00
Rafael dos Santos Silva	04bc402aae	FEATURE: Setting to control per post embeddings (#439 ) * FEATURE: Setting to control per post embeddings	2024-01-23 22:09:27 -03:00
Rafael dos Santos Silva	d4e23e0df6	FIX: Don't try to generate embeddings of posts in deleted topics (#433 )	2024-01-18 16:10:25 -03:00
Roman Rizzi	f9d7d7f5f0	DEV: AI bot migration to the Llm pattern. (#343 ) * DEV: AI bot migration to the Llm pattern. We added tool and conversation context support to the Llm service in discourse-ai#366, meaning we met all the conditions to migrate this module. This PR migrates to the new pattern, meaning adding a new bot now requires minimal effort as long as the service supports it. On top of this, we introduce the concept of a "Playground" to separate the PM-specific bits from the completion, allowing us to use the bot in other contexts like chat in the future. Commands are called tools, and we simplified all the placeholder logic to perform updates in a single place, making the flow more one-wayish. * Followup fixes based on testing * Cleanup unused inference code * FIX: text-based tools could be in the middle of a sentence * GPT-4-turbo support * Use new LLM API	2024-01-04 10:44:07 -03:00
Rafael dos Santos Silva	140359c2ef	FEATURE: Per post embeddings (#387 )	2023-12-29 12:28:45 -03:00
Keegan George	6aaf1f002e	FEATURE: Add streaming to post AI helper's explain option (#344 ) Co-authored-by: Rafael dos Santos Silva <xfalcox@gmail.com> Co-authored-by: Roman Rizzi <roman@discourse.org>	2023-12-12 09:28:39 -08:00
Sam	6ddc17fd61	DEV: port directory structure to Zeitwerk (#319 ) Previous to this change we relied on explicit loading for a files in Discourse AI. This had a few downsides: - Busywork whenever you add a file (an extra require relative) - We were not keeping to conventions internally ... some places were OpenAI others are OpenAi - Autoloader did not work which lead to lots of full application broken reloads when developing. This moves all of DiscourseAI into a Zeitwerk compatible structure. It also leaves some minimal amount of manual loading (automation - which is loading into an existing namespace that may or may not be there) To avoid needing /lib/discourse_ai/... we mount a namespace thus we are able to keep /lib pointed at ::DiscourseAi Various files were renamed to get around zeitwerk rules and minimize usage of custom inflections Though we can get custom inflections to work it is not worth it, will require a Discourse core patch which means we create a hard dependency.	2023-11-29 15:17:46 +11:00
Roman Rizzi	ef6c785aca	DEV: Move jobs undear each module lib directory	2023-02-23 14:09:52 -03:00
Roman Rizzi	1afa274b99	DEV: Reorganize files and add an entry point for each module	2023-02-23 12:25:00 -03:00
Rafael dos Santos Silva	6cf411ec90	add toxicity and sentiment modules	2023-02-22 20:46:53 -03:00

47 Commits