discourse-ai

mirror of https://github.com/discourse/discourse-ai.git synced 2025-02-18 17:34:52 +00:00

Author	SHA1	Message	Date
Roman Rizzi	e52045ebdc	DEV: Robust check for embeddings enabled (#1116 )	2025-02-06 12:18:55 -03:00
Roman Rizzi	e2e753d73c	FEATURE: Formalize support for matryoshka dimensions. (#1083 ) We have a flag to signal we are shortening the embeddings of a model. Only used in Open AI's text-embedding-3-*, but we plan to use it for other services.	2025-01-22 11:26:46 -03:00
Roman Rizzi	a5e5ae72a8	FIX: Open AI embedding shortening is only available for some models (#1080 )	2025-01-21 17:50:40 -03:00
Roman Rizzi	3b66fb3e87	FIX: Restore the accidentally deleted query prefix. (#1079 ) Additionally, we add a prefix for embedding generation. Both are stored in the definitions table.	2025-01-21 14:10:31 -03:00
Roman Rizzi	f5cf1019fb	FEATURE: configurable embeddings (#1049 ) * Use AR model for embeddings features * endpoints * Embeddings CRUD UI * Add presets. Hide a couple more settings * system specs * Seed embedding definition from old settings * Generate search bit index on the fly. cleanup orphaned data * support for seeded models * Fix run test for new embedding * fix selected model not set correctly	2025-01-21 12:23:19 -03:00
Roman Rizzi	534b0df391	REFACTOR: Separation of concerns for embedding generation. (#1027 ) In a previous refactor, we moved the responsibility of querying and storing embeddings into the `Schema` class. Now, it's time for embedding generation. The motivation behind these changes is to isolate vector characteristics in simple objects to later replace them with a DB-backed version, similar to what we did with LLM configs.	2024-12-16 09:55:39 -03:00
Roman Rizzi	eae527f99d	REFACTOR: A Simpler way of interacting with embeddings tables. (#1023 ) * REFACTOR: A Simpler way of interacting with embeddings' tables. This change adds a new abstraction called `Schema`, which acts as a repository that supports the same DB features `VectorRepresentation::Base` has, with the exception that removes the need to have duplicated methods per embeddings table. It is also a bit more flexible when performing a similarity search because you can pass it a block that gives you access to the builder, allowing you to add multiple joins/where conditions.	2024-12-13 10:15:21 -03:00
Roman Rizzi	6da35d8e66	FIX: Gemini inference client was missing #instance (#1019 )	2024-12-10 15:42:31 -03:00
Roman Rizzi	b32b1cf241	FIX: Add a digest check to avoid repeteadly generating embeddings (bulk) (#1001 )	2024-12-04 17:47:28 -03:00
Sam	0cb2c413ba	FEATURE: exclude muted categories from category suggester (#979 ) The logic here is that users do not particularly care about topics in the category so we can exclude them from tag and category suggestions	2024-11-29 12:17:28 +11:00
Roman Rizzi	ef07fcb308	FIX: Skip records without content to classify (#960 )	2024-11-26 15:54:20 -03:00
Roman Rizzi	ddf2bf7034	DEV: Backfill embeddings concurrently. (#941 ) We are adding a new method for generating and storing embeddings in bulk, which relies on `Concurrent::Promises::Future`. Generating an embedding consists of three steps: Prepare text HTTP call to retrieve the vector Save to DB. Each one is independently executed on whatever thread the pool gives us. We are bringing a custom thread pool instead of the global executor since we want control over how many threads we spawn to limit concurrency. We also avoid firing thousands of HTTP requests when working with large batches.	2024-11-26 14:12:32 -03:00
Mark VanLandingham	52d90cf1bc	DEV: Add apply_modifier for SemanticTopicQuery topics list (#830 )	2024-10-10 12:13:16 -05:00
Roman Rizzi	8849caf136	DEV: Transition "Select model" settings to only use LlmModels (#675 ) We no longer support the "provider:model" format in the "ai_helper_model" and "ai_embeddings_semantic_search_hyde_model" settings. We'll migrate existing values and work with our new data-driven LLM configs from now on.	2024-06-19 18:01:35 -03:00
Loïc Guitaut	6ae4218a96	DEV: Fix new Rubocop offenses	2024-03-06 15:23:29 +01:00
Sam	ba3c3951cf	FIX: typo causing text_embedding_3_large to fail (#460 )	2024-02-05 11:16:36 +11:00
Roman Rizzi	fba9c1bf2c	UX: Re-introduce embedding settings validations (#457 ) * Revert "Revert "UX: Validate embeddings settings (#455)" (#456)" This reverts commit 392e2e8aef7d5b0d988b3c3bc5cc19f1d83c4491. * Resstore previous default	2024-02-01 16:54:09 -03:00
Roman Rizzi	392e2e8aef	Revert "UX: Validate embeddings settings (#455 )" (#456 ) This reverts commit 85fca89e011933a0479abaf4bf0945983fb948b8.	2024-02-01 14:06:51 -03:00
Roman Rizzi	85fca89e01	UX: Validate embeddings settings (#455 )	2024-02-01 13:05:38 -03:00
Sam	dcafc8032f	FIX: improve embedding generation (#452 ) 1. on failure we were queuing a job to generate embeddings, it had the wrong params. This is both fixed and covered in a test. 2. backfill embedding in the order of bumped_at, so newest content is embedded first, cover with a test 3. add a safeguard for hidden site setting that only allows batches of 50k in an embedding job run Previously old embeddings were updated in a random order, this changes it so we update in a consistent order	2024-01-31 10:38:47 -03:00
Roman Rizzi	0634b85a81	UX: Validations to LLM-backed features (except AI Bot) (#436 ) * UX: Validations to Llm-backed features (except AI Bot) This change is part of an ongoing effort to prevent enabling a broken feature due to lack of configuration. We also want to explicit which provider we are going to use. For example, Claude models are available through AWS Bedrock and Anthropic, but the configuration differs. Validations are: * You must choose a model before enabling the feature. * You must turn off the feature before setting the model to blank. * You must configure each model settings before being able to select it. * Add provider name to summarization options * vLLM can technically support same models as HF * Check we can talk to the selected model * Check for Bedrock instead of anthropic as a site could have both creds setup	2024-01-29 16:04:25 -03:00
Rafael dos Santos Silva	04bc402aae	FEATURE: Setting to control per post embeddings (#439 ) * FEATURE: Setting to control per post embeddings	2024-01-23 22:09:27 -03:00
Jarek Radosz	6b8a57d957	DEV: Update linting (#423 ) Co-authored-by: Keegan George <kgeorge13@gmail.com>	2024-01-13 00:28:06 +01:00
Rafael dos Santos Silva	3be76ebd7a	FEATURE: Move the default embeddings model to bge-large-en (#417 )	2024-01-11 14:16:25 -03:00
Rafael dos Santos Silva	140359c2ef	FEATURE: Per post embeddings (#387 )	2023-12-29 12:28:45 -03:00
Krzysztof Kotlarek	6de9b9c274	DEV: Replace deprecated min_trust_to_create_post (#356 ) In https://github.com/discourse/discourse/pull/24740, `min_trust_to_create_topic` site setting was replaced by `create_topic_allowed_groups`. This PR replaces the former, deprecated one, with the latter.	2023-12-14 14:07:28 +11:00
Martin Brennan	24370a9ca6	Revert "FIX: Use Guardian.basic_user instead of new (anon) (#332 )" (#337 ) This reverts commit a3a1285dc5f839b7b83649113b0b2c700e0fde77. c.f. https://github.com/discourse/discourse/pull/24742	2023-12-06 16:26:43 +10:00
Martin Brennan	a3a1285dc5	FIX: Use Guardian.basic_user instead of new (anon) (#332 ) c.f. de983796e1b66aa2ab039a4fb6e32cec8a65a098 There will soon be additional login_required checks for Guardian, and the intent of many checks by automated systems is better fulfilled by using BasicUser, which simulates a logged in TL0 forum user, rather than an anon user.	2023-12-06 12:01:41 +10:00
Sam	6ddc17fd61	DEV: port directory structure to Zeitwerk (#319 ) Previous to this change we relied on explicit loading for a files in Discourse AI. This had a few downsides: - Busywork whenever you add a file (an extra require relative) - We were not keeping to conventions internally ... some places were OpenAI others are OpenAi - Autoloader did not work which lead to lots of full application broken reloads when developing. This moves all of DiscourseAI into a Zeitwerk compatible structure. It also leaves some minimal amount of manual loading (automation - which is loading into an existing namespace that may or may not be there) To avoid needing /lib/discourse_ai/... we mount a namespace thus we are able to keep /lib pointed at ::DiscourseAi Various files were renamed to get around zeitwerk rules and minimize usage of custom inflections Though we can get custom inflections to work it is not worth it, will require a Discourse core patch which means we create a hard dependency.	2023-11-29 15:17:46 +11:00
Roman Rizzi	3064d4c288	REFACTOR: Summarization and HyDE now use an LLM abstraction. (#297 ) * DEV: One LLM abstraction to rule them all * REFACTOR: HyDE search uses new LLM abstraction * REFACTOR: Summarization uses the LLM abstraction * Updated documentation and made small fixes. Remove Bedrock claude-2 restriction	2023-11-23 12:58:54 -03:00
Sam	98c89953d3	FEATURE: remember previously selected persona (#299 ) People tend to keep to 1 persona when working with the bot, this adds local browser memory for the last persona you interacted with so you do not need to select it over and over again. This is per browser, not per user memory. Also... clean up tests so they do not need to require stubs which were breaking the build --------- Co-authored-by: Martin Brennan <martin@discourse.org>	2023-11-21 17:02:27 +10:00
Sam	5b5edb22c6	FEATURE: UI to update ai personas on admin page (#290 ) Introduces a UI to manage customizable personas (admin only feature) Part of the change was some extensive internal refactoring: - AIBot now has a persona set in the constructor, once set it never changes - Command now takes in bot as a constructor param, so it has the correct persona and is not generating AIBot objects on the fly - Added a .prettierignore file, due to the way ALE is configured in nvim it is a pre-req for prettier to work - Adds a bunch of validations on the AIPersona model, system personas (artist/creative etc...) are all seeded. We now ensure - name uniqueness, and only allow certain properties to be touched for system personas. - (JS note) the client side design takes advantage of nested routes, the parent route for personas gets all the personas via this.store.findAll("ai-persona") then child routes simply reach into this model to find a particular persona. - (JS note) data is sideloaded into the ai-persona model the meta property supplied from the controller, resultSetMeta - This removes ai_bot_enabled_personas and ai_bot_enabled_chat_commands, both should be controlled from the UI on a per persona basis - Fixes a long standing bug in token accounting ... we were doing to_json.length instead of to_json.to_s.length - Amended it so {commands} are always inserted at the end unconditionally, no need to add it to the template of the system message as it just confuses things - Adds a concept of required_commands to stock personas, these are commands that must be configured for this stock persona to show up. - Refactored tests so we stop requiring inference_stubs, it was very confusing to need it, added to plugin.rb for now which at least is clearer - Migrates the persona selector to gjs --------- Co-authored-by: Joffrey JAFFEUX <j.jaffeux@gmail.com> Co-authored-by: Martin Brennan <martin@discourse.org>	2023-11-21 16:56:43 +11:00
Roman Rizzi	0828254d61	FIX: Generate embeddings job was broken (#211 ) * FIX: Use correct methods to generate embeddings * FIX: Generate embeddings job was broken	2023-09-07 11:54:43 -03:00
Roman Rizzi	13d63f1f30	FIX: filter allowed categories from semantic search results (#206 )	2023-09-06 10:00:20 -03:00
Rafael dos Santos Silva	4b42c09814	FEATURE: Tweak HyDE prompts for better grounding in forum subject and limit response size (#200 ) * FEATURE: Tweak HyDE prompts for better grounding in forum subject and limit response size * fix test * lint	2023-09-05 16:11:07 -03:00
Rafael dos Santos Silva	2c0f535bab	FEATURE: HyDE-powered semantic search. (#136 ) * FEATURE: HyDE-powered semantic search. It relies on the new outlet added on discourse/discourse#23390 to display semantic search results in an unobtrusive way. We'll use a HyDE-backed approach for semantic search, which consists on generating an hypothetical document from a given keywords, which gets transformed into a vector and used in a asymmetric similarity topic search. This PR also reorganizes the internals to have less moving parts, maintaining one hierarchy of DAOish classes for vector-related operations like transformations and querying. Completions and vectors created by HyDE will remain cached on Redis for now, but we could later use Postgres instead. * Missing translation and rate limiting --------- Co-authored-by: Roman Rizzi <rizziromanalejandro@gmail.com>	2023-09-05 11:08:23 -03:00
Rafael dos Santos Silva	0738f67fa4	FIX: Fix embeddings truncation strategy (#139 )	2023-08-16 15:09:41 -03:00
Rafael dos Santos Silva	8318c4374c	FIX: Remove muted from Similar list (#127 ) * FIX: Remove muted from Similar list	2023-08-08 15:44:10 -03:00
Roman Rizzi	58b96eda6c	REFACTOR: Build related topics using TopicQuery. (#124 ) TopicQuery already provides a lot of safeguards and options for filtering topic, and enforcing permissions. It makes sense to rely on it as other plugins like discourse-assign do. As a bonus, we now have access to the current_user while serializing these topics, so users will see things like unread posts count just like we do for the lists.	2023-08-02 16:58:09 -03:00
Rafael dos Santos Silva	5e3f4e1b78	FEATURE: Embeddings to main db (#99 ) * FEATURE: Embeddings to main db This commit moves our embeddings store from an external configurable PostgreSQL instance back into the main database. This is done to simplify the setup. There is a migration that will try to import the external embeddings into the main DB if it is configured and there are rows. It removes support from embeddings models that aren't all_mpnet_base_v2 or OpenAI text_embedding_ada_002. However it will now be easier to add new models. It also now takes into account: - topic title - topic category - topic tags - replies (as much as the model allows) We introduce an interface so we can eventually support multiple strategies for handling long topics. This PR severely damages the semantic search performance, but this is a temporary until we can get adapt HyDE to make semantic search use the same embeddings we have for semantic related with good performance. Here we also have some ground work to add post level embeddings, but this will be added in a future PR. Please note that this PR will also block Discourse from booting / updating if this plugin is installed and the pgvector extension isn't available on the PostgreSQL instance Discourse uses.	2023-07-13 12:41:36 -03:00
Sam	b82fc1e692	FIX: ensure we only attempt embedding once every 15 minutes (#76 ) This also heavily reduced log noise and ensures our exception handling is more surgical.	2023-05-23 10:43:24 +10:00
Rafael dos Santos Silva	e5537d4c77	FEATURE: Allow excluding closed topics from semantic related (#55 )	2023-05-09 15:30:50 -03:00
Roman Rizzi	4e05763a99	FEATURE: Semantic assymetric full-page search (#34 ) Depends on discourse/discourse#20915 Hooks to the full-page-search component using an experimental API and performs an assymetric similarity search using our embeddings database.	2023-03-31 15:29:56 -03:00
Sam	6543c50758	FIX: stop returning self as a candidate for related topics (#31 )	2023-03-31 11:04:17 +10:00
Sam	0d80d9ec49	FEATURE: allow limiting results in related topics section (#30 ) Also: - Normalizes behavior between logged in and anon, we only show related topics in the related topic section - Renames "suggested" to "related" given this only exists in related section - Adds a spec section to ensure anon does not regress - Adds `ai_embeddings_semantic_related_topics` to limit related topics Renamed settings: ai_embeddings_semantic_suggested_model -> ai_embeddings_semantic_related_model ai_embeddings_semantic_suggested_topics_enabled -> ai_embeddings_semantic_related_topics_enabled Plugins is still in an experimental phase and not much is overidden hence avoiding adding site setting migrations. Co-authored-by: Krzysztof Kotlarek <kotlarek.krzysztof@gmail.com>	2023-03-31 11:04:34 +11:00
Sam	1d097b9d82	FEATURE: attempt to include related topics above suggested (#28 ) Allows related topics to show up for logged on users - Introduces a new "Related Topics" block above suggested when related topics exist - Renames `ai_embeddings_semantic_suggested_topics_anons_enabled` -> `ai_embeddings_semantic_suggested_topics_enabled` (given it is only deployed on 1 site not bothering with a migration) - Adds an integration test to ensure data arrives correctly on the client	2023-03-31 09:07:22 +11:00
Rafael dos Santos Silva	45950f1bb4	FIX: Only show public visible topics as suggested for anons (#27 ) * FIX: Only show public visible topics as suggested for anons * DEV: Add tests for embeddings * Update spec/lib/modules/embeddings/semantic_suggested_spec.rb Co-authored-by: Bianca Nenciu <nbianca@users.noreply.github.com> * Update spec/lib/modules/embeddings/semantic_suggested_spec.rb Co-authored-by: Bianca Nenciu <nbianca@users.noreply.github.com> * move to top --------- Co-authored-by: Bianca Nenciu <nbianca@users.noreply.github.com>	2023-03-23 17:28:01 -03:00

47 Commits