discourse-ai

Commit Graph

Author	SHA1	Message	Date
Sam	b7a96e3bcb	FIX: avoid all bot feedback loops (#507 ) We need to ensure that under no circumstances feedback loops between bots will emerge cause this can eat up a lot of tokens	2024-03-05 10:02:49 +11:00
Sam	77cf9e2cff	FIX: system persona non English save, missing bot pms - FIX: only update system attributes when updating system persona - FIX: update participant count by hand so bot messages show in inbox Co-authored-by: Joffrey JAFFEUX <j.jaffeux@gmail.com>	2024-03-04 09:56:59 +11:00
Sam	c02794cf2e	FIX: support multiple tool calls (#502 ) * FIX: support multiple tool calls Prior to this change we had a hard limit of 1 tool call per llm round trip. This meant you could not google multiple things at once or perform searches across two tools. Also: - Hint when Google stops working - Log topic_id / post_id when performing completions * Also track id for title	2024-03-02 07:53:21 +11:00
Sam	59bab2bba3	FIX: stream messages when directly PMing a persona (#500 ) previous to this fix we did not consider personas a bot in the front end	2024-03-01 07:53:42 +11:00
Sam	9fb1430e40	FIX: support spaces within arguments for Open AI (#499 ) Previous to this fix if a tool call ever streamed a SPACE alone, we would eat it and ignore it, breaking params Also fixes some tests to ensure they are actually called :)	2024-02-29 12:47:34 +11:00
Rafael dos Santos Silva	1b72a00d2c	FEATURE: Option for AI triage to send a post to the review queue (#498 ) Option for AI triage to send a post to the review queue	2024-02-29 12:33:28 +11:00
Sam	484fd1435b	DEV: improve internal design of ai persona and bug fix (#495 ) * DEV: improve internal design of ai persona and bug fix - Fixes bug where OpenAI could not describe images - Fixes bug where mentionable personas could not be mentioned unless overarching bot was enabled - Improves internal design of playground and bot to allow better for non "bot" users - Allow PMs directly to persona users (previously bot user would also have to be in PM) - Simplify internal code Co-authored-by: Martin Brennan <martin@discourse.org>	2024-02-28 16:46:32 +11:00
Rafael dos Santos Silva	a1f1067f69	FIX: Lower truncation size for Gemini Embeddings (#493 )	2024-02-28 08:52:53 +11:00
Sam	d036f3fb8e	FEATURE: AI helper support in non English languages (#489 ) * FEATURE: AI helper support in non English languages This attempts some prompt engineering to coerce AI helper to answer in the appropriate language. Note mileage will vary, in testing GPT-4 produces the best results GPT-3.5 can return OKish results. * Extend non english support for GPT-4V image caption * Update db/fixtures/ai_helper/603_completion_prompts.rb --------- Co-authored-by: Rafael Silva <xfalcox@gmail.com>	2024-02-27 16:31:51 -03:00
Sam	aabff87501	FIX: image generation in gemini was broken (#490 ) We need to inject blank model answers after tool calls if absent otherwise model will reject it.	2024-02-27 18:24:30 +11:00
Roman Rizzi	94ba0dadc2	SECURITY: Place a SSRF protection when calling services from the plugin. (#485 ) The Faraday adapter and `FinalDestionation::HTTP` will protect us from admin-initiated SSRF attacks when interacting with the external services powering this plugin features.:	2024-02-21 17:14:50 -03:00
Sam	becbe01f68	FIX: unable to share conversations with persona user (#479 ) Persona users are still bots, but we were not properly accounting for it and share icon was not showing up. This depends on a core change that adds .topic to transformed posts	2024-02-20 16:16:23 +11:00
Keegan George	a9b2d6a30a	FEATURE: AI image caption (#470 ) This PR adds a new feature where you can generate captions for images in the composer using AI. --------- Co-authored-by: Rafael Silva <xfalcox@gmail.com>	2024-02-19 14:56:28 -03:00
Sam	1f74a77e17	DEV: correct flaky spec (#475 ) We were not properly expiring prompt cache	2024-02-19 15:21:55 +11:00
Sam	0fb87b00e2	FEATURE: new Discourse Helper persona (#473 ) This persona searches Discourse Meta for help with Discourse and points users at relevant posts. It is somewhat similar to using "Forum Helper" on meta, with the notable difference that we can not lean on semantic search so using some prompt engineering we try to keep it simple.	2024-02-19 14:52:12 +11:00
Krzysztof Kotlarek	dd6b073fc3	DEV: Make more group-based settings client: false (#474 ) Affects the following settings: ai_toxicity_groups_bypass ai_helper_allowed_groups ai_helper_custom_prompts_allowed_groups post_ai_helper_allowed_groups This turns off client: true for these group-based settings, because there is no guarantee that the current user gets all their group memberships serialized to the client. Better to check server-side first.	2024-02-19 13:26:24 +11:00
Keegan George	d66915ecc1	DEV: Make prompts available on `CurrentUserSerializer` (#472 )	2024-02-16 10:57:14 -08:00
Sam	3a8d95f6b2	FEATURE: mentionable personas and random picker tool, context limits (#466 ) 1. Personas are now optionally mentionable, meaning that you can mention them either from public topics or PMs - Mentioning from PMs helps "switch" persona mid conversation, meaning if you want to look up sites setting you can invoke the site setting bot, or if you want to generate an image you can invoke dall e - Mentioning outside of PMs allows you to inject a bot reply in a topic trivially - We also add the support for max_context_posts this allow you to limit the amount of context you feed in, which can help control costs 2. Add support for a "random picker" tool that can be used to pick random numbers 3. Clean up routing ai_personas -> ai-personas 4. Add Max Context Posts so users can control how much history a persona can consume (this is important for mentionable personas) Co-authored-by: Martin Brennan <martin@discourse.org>	2024-02-15 16:37:59 +11:00
Rafael dos Santos Silva	59fbbb156b	DEV: Make indexing less frequent when related topics is disabled (#468 )	2024-02-09 16:08:54 -03:00
Rafael dos Santos Silva	0dba6623a0	FIX: Better AI chat thread titles (#467 ) * FIX: Better AI chat thread titles - Fix quote removal when multi-line - Use XML tags for better LLM output parsing - Use stop_sequences for faster and less wasteful LLM calls - Adds truncation as the last line of defense	2024-02-09 14:49:28 -03:00
Rafael dos Santos Silva	bccb7efdd6	FIX: Use a dedicated prompt for thread titles (#464 )	2024-02-07 15:05:50 -03:00
Sam	ba3c3951cf	FIX: typo causing text_embedding_3_large to fail (#460 )	2024-02-05 11:16:36 +11:00
Sam	a3c827efcc	FEATURE: allow personas to supply top_p and temperature params (#459 ) * FEATURE: allow personas to supply top_p and temperature params Code assistance generally are more focused at a lower temperature This amends it so SQL Helper runs at 0.2 temperature vs the more common default across LLMs of 1.0. Reduced temperature leads to more focused, concise and predictable answers for the SQL Helper * fix tests * This is not perfect, but far better than what we do today Instead of fishing for 1. Draft sequence 2. Draft body We skip (2), this means the composer "only" needs 1 http request to open, we also want to eliminate (1) but it is a bit of a trickier core change, may figure out how to pull it off (defer it to first draft save) Value of bot drafts < value of opening bot conversations really fast	2024-02-03 07:09:34 +11:00
Roman Rizzi	fba9c1bf2c	UX: Re-introduce embedding settings validations (#457 ) * Revert "Revert "UX: Validate embeddings settings (#455)" (#456)" This reverts commit `392e2e8aef`. * Resstore previous default	2024-02-01 16:54:09 -03:00
Roman Rizzi	392e2e8aef	Revert "UX: Validate embeddings settings (#455 )" (#456 ) This reverts commit `85fca89e01`.	2024-02-01 14:06:51 -03:00
Roman Rizzi	85fca89e01	UX: Validate embeddings settings (#455 )	2024-02-01 13:05:38 -03:00
Sam	cec4251b00	DEV: improve error bedrock error messages (#454 ) When bedrock rate limits it returns a 200 BUT also returns a JSON document with the error. Previously we had no special case here so we complained about nil New code properly logs the problem	2024-02-01 08:01:07 -03:00
Sam	dcafc8032f	FIX: improve embedding generation (#452 ) 1. on failure we were queuing a job to generate embeddings, it had the wrong params. This is both fixed and covered in a test. 2. backfill embedding in the order of bumped_at, so newest content is embedded first, cover with a test 3. add a safeguard for hidden site setting that only allows batches of 50k in an embedding job run Previously old embeddings were updated in a random order, this changes it so we update in a consistent order	2024-01-31 10:38:47 -03:00
Sam	abcf5ea94a	FEATURE: fine tune llm report to follow instructions more closely (#451 ) - Allow users to supply top_p and temperature values, which means people can fine tune randomness - Fix bad localization string - Fix bad remapping of max tokens in gemini - Add support for top_p as a general param to llms - Amend system prompt so persona stops treating a user as an adversary	2024-01-31 09:58:25 +11:00
Rafael dos Santos Silva	b41c5cc31c	FIX: Add table name to remove ambiguous column reference in SQL (#449 )	2024-01-30 15:50:26 -03:00
Sam	ab7e9e31aa	FEATURE: allow excluding tags and categories from LLM report (#447 ) Also - Better diagnostics, output model being used - Prompt LLM that true content is being injected in <context> tag	2024-01-30 15:55:05 +11:00
Roman Rizzi	bae71eb047	FIX: Include provider in automation models (#446 )	2024-01-29 18:07:29 -03:00
Roman Rizzi	0634b85a81	UX: Validations to LLM-backed features (except AI Bot) (#436 ) * UX: Validations to Llm-backed features (except AI Bot) This change is part of an ongoing effort to prevent enabling a broken feature due to lack of configuration. We also want to explicit which provider we are going to use. For example, Claude models are available through AWS Bedrock and Anthropic, but the configuration differs. Validations are: * You must choose a model before enabling the feature. * You must turn off the feature before setting the model to blank. * You must configure each model settings before being able to select it. * Add provider name to summarization options * vLLM can technically support same models as HF * Check we can talk to the selected model * Check for Bedrock instead of anthropic as a site could have both creds setup	2024-01-29 16:04:25 -03:00
Sam	b2b01185f2	FEATURE: add support for new OpenAI embedding models (#445 ) * FEATURE: add support for new OpenAI embedding models This adds support for just released text_embedding_3_small and large Note, we have not yet implemented truncation support which is a new API feature. (triggered using dimensions) * Tiny side fix, recalc bots when ai is enabled or disabled * FIX: downsample to 2000 items per vector which is a pgvector limitation	2024-01-29 13:24:30 -03:00
Sam	092da860e2	FEATURE: support gpt-4-0125 which was just released (#443 ) The new model has better performance and is always preferable to the old one which has unicode issues during function calls.	2024-01-26 09:08:02 +11:00
Roman Rizzi	b461ebc4ca	FIX: typo in Automation::AVAILABLE_MODELS (#442 )	2024-01-25 11:56:28 -03:00
Rafael dos Santos Silva	fa6bc7f409	FIX: Automatic embeddings index could fail if it existed in the backup schema (#441 )	2024-01-24 15:57:26 -03:00
Rafael dos Santos Silva	16d666fe69	FIX: Misconfigured OpenAI API for embeddings shouldn't spam logs (#440 )	2024-01-24 15:57:18 -03:00
Rafael dos Santos Silva	04bc402aae	FEATURE: Setting to control per post embeddings (#439 ) * FEATURE: Setting to control per post embeddings	2024-01-23 22:09:27 -03:00
Jarek Radosz	5802cd1a0c	DEV: Fix various typos (#434 )	2024-01-19 12:51:26 +01:00
Rafael dos Santos Silva	c70f43f130	FIX: Truncate content for sentiment/toxicity classification (#431 )	2024-01-17 15:17:58 -03:00
Roman Rizzi	5bdf3dc1f4	DEV: Stop using shared_examples for endpoint specs (#430 )	2024-01-17 15:08:49 -03:00
Sam	370074ef21	FIX: always ensure `#generate` gets a valid input (#427 ) We were not validating input for generate leading to 2 tests not failing correctly despite functionality being broken. This ensures that input is validated,and in turn fixes the broken specs	2024-01-16 15:21:58 +11:00
Sam	05d8b021f1	FIX: scrub invalid prompts when truncating (#426 ) When you trim a prompt we never want to have a state where there is a "tool" reply without a corresponding tool call, it makes no sense Also - GPT-4-Turbo is 128k, fix that - Claude was not preserving username in prompt - We were throwing away unicode usernames instead of adding to message	2024-01-16 13:48:00 +11:00
Roman Rizzi	ff4da6ace8	FIX: Clean unicode usernames when adding messages through prompt's contrstuctor (#425 )	2024-01-15 12:01:40 -03:00
Sam	825f01cfb2	FEATURE: even smoother streaming (#420 ) Account properly for function calls, don't stream through <details> blocks - Rush cooked content back to client - Wait longer (up to 60 seconds) before giving up on streaming - Clean up message bus channels so we don't have leftover data - Make ai streamer much more reusable and much easier to read - If buffer grows quickly, rush update so you are not artificially waiting - Refine prompt interface - Fix lost system message when prompt gets long	2024-01-15 18:51:14 +11:00
Jarek Radosz	6b8a57d957	DEV: Update linting (#423 ) Co-authored-by: Keegan George <kgeorge13@gmail.com>	2024-01-13 00:28:06 +01:00
Roman Rizzi	04eae76f68	REFACTOR: Represent generic prompts with an Object. (#416 ) * REFACTOR: Represent generic prompts with an Object. * Adds a bit more validation for clarity * Rewrite bot title prompt and fix quirk handling --------- Co-authored-by: Sam Saffron <sam.saffron@gmail.com>	2024-01-12 14:36:44 -03:00
Rafael dos Santos Silva	705ef986b4	FIX: Set ivfflat.probes using topic count, not post count (#421 ) Fixes a regression from `140359c` which caused we to set this globally based on post count, rendering the cost of an index scan on the topics table too high and making the planner, correctly, not use the index anymore. Hopefully https://github.com/pgvector/pgvector/issues/235 lands soon.	2024-01-12 11:20:23 -03:00
Sam	8df966e9c5	FEATURE: smooth streaming of AI responses on the client (#413 ) This PR introduces 3 things: 1. Fake bot that can be used on local so you can test LLMs, to enable on dev use: SiteSetting.ai_bot_enabled_chat_bots = "fake" 2. More elegant smooth streaming of progress on LLM completion This leans on JavaScript to buffer and trickle llm results through. It also amends it so the progress dot is much more consistently rendered 3. It fixes the Claude dialect Claude needs newlines exactly at the right spot, amended so it is happy --------- Co-authored-by: Martin Brennan <martin@discourse.org>	2024-01-11 15:56:40 +11:00
Martin Brennan	37b957dbbb	DEV: Fix SemanticRelated module load error (#419 ) Followup `2636efcd1b`, whenever ruby code was changed locally this would break module loading, giving an "uninitialized constant DiscourseAi::Embeddings::EntryPoint::SemanticRelated" error.	2024-01-11 13:52:50 +10:00
Rafael dos Santos Silva	8fcba12fae	FEATURE: Support for SRV records for Discourse services (#414 ) This allows admins to configure services with multiple backends using DNS SRV records. This PR also adds support for shared secret auth via headers for TEI and vLLM endpoints, so they are inline with the other ones.	2024-01-10 19:23:07 -03:00
Roman Rizzi	abde82c1f3	FIX: Use claude-2.1 to enable system prompts (#411 )	2024-01-09 14:10:20 -03:00
Sam	05f7808057	FEATURE: more elegant progress (#409 ) Previous to this change it was very hard to tell if completion was stuck or not. This introduces a "dot" that follows the completion and starts flashing after 5 seconds.	2024-01-09 09:20:28 -03:00
Sam	b0a0cbe3ca	FIX: improve bot behavior (#408 ) * FIX: improve bot behavior - Provide more information to Gemini context post function execution - Use system prompts for Claude (fixes Dall E) - Ensure Assistant is properly separated - Teach Claude to return arrays in JSON vs XML Also refactors tests so we do not copy tool preamble everywhere * System msg is claude-2 only. fix typo --------- Co-authored-by: Roman Rizzi <rizziromanalejandro@gmail.com>	2024-01-08 10:28:03 -03:00
Roman Rizzi	6124f910c1	FIX: Bring back Azure support. (#407 ) We thought Azure's latest API version didn't have tool support yet, but I didn't understand it was complaining about a required field in the tool call message.	2024-01-05 17:08:10 -03:00
Sam	17cc09ec9c	FIX: don't include <details> in context (#406 ) * FIX: don't include <details> in context We need to be careful adding <details> into context of conversations it can cause LLMs to hallucinate results * Fix Gemini multi-turn ctx flattening --------- Co-authored-by: Roman Rizzi <rizziromanalejandro@gmail.com>	2024-01-05 15:21:14 -03:00
Keegan George	7201d482d5	FEATURE: Add DallE support to AI helper's illustrate post (#404 )	2024-01-05 09:03:23 -08:00
Rafael dos Santos Silva	23b2809638	FEATURE: Generate proper embeddings for posts/topics with embedded content (#401 )	2024-01-05 10:27:45 -03:00
Rafael dos Santos Silva	6fc1c9f7a6	FEATURE: Try to automatically handle larger embedding indexes (#403 ) * FEATURE: Try to automatically handle larger embedding indexes * linteeeeeeeer	2024-01-05 09:56:28 -03:00
Sam	dd42a4e47b	FIX: array arguments not parsed correctly (#405 ) DALL E command accepts an Array as a tool argument, this was not parsed correctly by the invoker leading to errors generating images with DALL E Side quest ... don't use update! it calls validations and will now fail due to email validation	2024-01-05 14:39:32 +11:00
Roman Rizzi	971e03bdf2	FEATURE: AI Bot Gemini support. (#402 ) It also corrects the syntax around tool support, which was wrong. Gemini doesn't want us to include messages about previous tool invocations, so I had to shuffle around some code to send the response it generated from those invocations instead. For this, I created the "multi_turn" context, which bundles all the context involved in the interaction.	2024-01-04 18:15:34 -03:00
Roman Rizzi	aa56baad37	FEATURE: Add Mixtral support for AI Bot (#396 )	2024-01-04 12:22:43 -03:00
Roman Rizzi	e6422c542e	FIX: Tools::DbSchema's tables parameter is a string (#400 )	2024-01-04 11:50:26 -03:00
Roman Rizzi	f9d7d7f5f0	DEV: AI bot migration to the Llm pattern. (#343 ) * DEV: AI bot migration to the Llm pattern. We added tool and conversation context support to the Llm service in discourse-ai#366, meaning we met all the conditions to migrate this module. This PR migrates to the new pattern, meaning adding a new bot now requires minimal effort as long as the service supports it. On top of this, we introduce the concept of a "Playground" to separate the PM-specific bits from the completion, allowing us to use the bot in other contexts like chat in the future. Commands are called tools, and we simplified all the placeholder logic to perform updates in a single place, making the flow more one-wayish. * Followup fixes based on testing * Cleanup unused inference code * FIX: text-based tools could be in the middle of a sentence * GPT-4-turbo support * Use new LLM API	2024-01-04 10:44:07 -03:00
Sam	03fc94684b	FIX: AI helper not working correctly with mixtral (#399 ) * FIX: AI helper not working correctly with mixtral This PR introduces a new function on the generic llm called #generate This will replace the implementation of completion! #generate introduces a new way to pass temperature, max_tokens and stop_sequences Then LLM implementers need to implement #normalize_model_params to ensure the generic names match the LLM specific endpoint This also adds temperature and stop_sequences to completion_prompts this allows for much more robust completion prompts * port everything over to #generate * Fix translation - On anthropic this no longer throws random "This is your translation:" - On mixtral this actually works * fix markdown table generation as well	2024-01-04 09:53:47 -03:00
Keegan George	0483e0bb88	UX: Add proper attribution to illustrate post images (#398 )	2024-01-03 13:01:19 -08:00
Keegan George	1a5985134a	FIX: Show illustrate post only if stability API key present (#395 )	2024-01-02 11:24:16 -08:00
Roman Rizzi	4182af230a	FIX: Correctly translate and read tools for Claude and Chat GPT. (#393 ) I tested against the live models for the AI bot migration. It ensures Open AI's tool syntax is correct and we can correctly read the replies. :	2024-01-02 11:21:13 -03:00
Rafael dos Santos Silva	cec9bb8910	FIX: Skip embeddings for blank content (#392 )	2023-12-29 14:59:08 -03:00
Rafael dos Santos Silva	c778592da4	FIX: Corner cases on post embedding and crawler related (#391 ) * FIX: Simplify markup for crawler related * FIX: Handle title-less topics when truncating for embeddings * fix	2023-12-29 14:05:02 -03:00
Rafael dos Santos Silva	140359c2ef	FEATURE: Per post embeddings (#387 )	2023-12-29 12:28:45 -03:00
Rafael dos Santos Silva	2636efcd1b	FEATURE: Render Related Topics for Crawlers (#386 )	2023-12-28 15:32:03 -03:00
Rafael dos Santos Silva	1287ef4428	FEATURE: Support for Gemini Embeddings (#382 )	2023-12-28 10:28:01 -03:00
Rafael dos Santos Silva	76f7940b55	Revert "FEATURE: User sentiment on profile summary page (#329 )" (#383 ) This reverts commit `71c5077228`.	2023-12-28 11:01:57 +11:00
Rafael dos Santos Silva	20cb15ab5f	FEATURE: Mixtral for summarization (#381 )	2023-12-26 17:50:02 -03:00
Rafael dos Santos Silva	3c27cbfb9a	FIX: Use vLLM if TGI is not configured for OSS LLM inference (#380 )	2023-12-26 17:18:08 -03:00
Rafael dos Santos Silva	5db7bf6e68	Mixtral (#376 ) Add both Mistral and Mixtral support. Also includes vLLM-openAI inference support. Co-authored-by: Roman Rizzi <rizziromanalejandro@gmail.com>	2023-12-26 14:49:55 -03:00
Sam	a5d240991f	FEATURE: allow sending AI based report to a topic (#377 ) This makes the reporting far more flexible cause it can target a far wider audience by pointing it at a topic in a secure category or an existing PM	2023-12-22 11:46:23 +11:00
Sam	37dd98c937	FIX: exclude non visible topics from report context (#375 ) Generally non visible topics are not that interesting, do not add this noise to the report context	2023-12-21 19:08:36 +11:00
Sam	af2e692761	FIX: under certain conditions we would get duplicate data from llm (#373 ) Previously endpoint/base would `+=` decoded_chunk to leftover This could lead to cases where the leftover buffer had duplicate previously processed data Fix ensures we properly skip previously decoded data.	2023-12-20 14:28:05 -03:00
Sam	8664771b7f	FIX: triage no longer working with claude (#369 )	2023-12-20 07:58:38 +11:00
Keegan George	5a84969c96	FIX: Illustrate post icon and translation not appearing correctly (#371 )	2023-12-19 12:55:43 -08:00
Keegan George	7b4710d5c9	FEATURE: Generate post illustrations (#367 )	2023-12-19 11:17:34 -08:00
Sam	529703b5ec	FEATURE: support sending AI report to an email address (#368 ) Support emailing the AI report to any arbitrary email	2023-12-19 17:51:49 +11:00
Sam	d0f54443ae	FEATURE: LLM based peroidical summary report (#357 ) Introduce a Discourse Automation based periodical report. Depends on Discourse Automation. Report works best with very large context language models such as GPT-4-Turbo and Claude 2. - Introduces final_insts to generic llm format, for claude to work best it is better to guide the last assistant message (we should add this to other spots as well) - Adds GPT-4 turbo support to generic llm interface	2023-12-19 12:04:15 +11:00
Roman Rizzi	e0bf6adb5b	DEV: Tool support for the LLM service. (#366 ) This PR adds tool support to available LLMs. We'll buffer tool invocations and return them instead of making users of this service parse the response. It also adds support for conversation context in the generic prompt. It includes bot messages, user messages, and tool invocations, which we'll trim to make sure it doesn't exceed the prompt limit, then translate them to the correct dialect. Finally, It adds some buffering when reading chunks to handle cases when streaming is extremely slow.:M	2023-12-18 18:06:01 -03:00
Roman Rizzi	203906be65	FIX: Bedrock was complaining input was too long (#365 )	2023-12-18 16:06:06 -03:00
Rafael dos Santos Silva	4d7ccdda2f	FEATURE: DNS SRV support for TEI (#363 )	2023-12-18 13:21:21 -03:00
Rafael dos Santos Silva	83744bf192	FEATURE: Support for Gemini in AiHelper / Search / Summarization (#358 )	2023-12-15 14:32:01 -03:00
Keegan George	408d9f68eb	FEATURE: Proofread with post AI helper (#359 )	2023-12-14 19:30:52 -08:00
Keegan George	74a7ac4a3d	FEATURE: Add custom prompts to post helper options (#355 ) * FEATURE: Add custom prompts to post helper options * 💄Make pretty * 💄Make pretty!	2023-12-14 13:47:20 -03:00
Roman Rizzi	031c2a6b46	Revert "FIX: Recover from Bedrock returning invalid base64 payloads during streaming (#352 )" (#353 ) This reverts commit `ef7d4cc509`.	2023-12-12 17:22:44 -03:00
Roman Rizzi	ef7d4cc509	FIX: Recover from Bedrock returning invalid base64 payloads during streaming (#352 )	2023-12-12 17:06:53 -03:00
Keegan George	6aaf1f002e	FEATURE: Add streaming to post AI helper's explain option (#344 ) Co-authored-by: Rafael dos Santos Silva <xfalcox@gmail.com> Co-authored-by: Roman Rizzi <roman@discourse.org>	2023-12-12 09:28:39 -08:00
Sam	605445831f	FEATURE: try including views/username/likes in search results (#349 ) This is somewhat experimental, but the context of likes/view/username can help the llm find out what content is more important or even common users that produce great content This inflates the amount of tokens somewhat, but given it is all numbers and search columns titles are only included once this is not severe	2023-12-12 12:22:28 +11:00
Roman Rizzi	2798e4c86d	FIX: Custom instructions where missing when generating custom prompt input (#348 )	2023-12-11 19:26:56 -03:00
Sam	a66b1042cc	FEATURE: scale up result count for search depending on model (#346 ) We were limiting to 20 results unconditionally cause we had to make sure search always fit in an 8k context window. Models such as GPT 3.5 Turbo (16k) and GPT 4 Turbo / Claude 2.1 (over 150k) allow us to return a lot more results. This means we have a much richer understanding cause context is far larger. This also allows a persona to tweak this number, in some cases admin may want to be conservative and save on tokens by limiting results This also tweaks the `limit` param which GPT-4 liked to set to tell model only to use it when it needs to (and describes default behavior)	2023-12-11 16:54:16 +11:00
Sam	3c9901d43a	FEATURE: implement GPT-4 turbo support (#345 ) Keep in mind: - GPT-4 is only going to be fully released next year - so this hardcodes preview model for now - Fixes streaming bugs which became a big problem with GPT-4 turbo - Adds Azure endpoing for turbo as well Co-authored-by: Martin Brennan <martin@discourse.org>	2023-12-11 14:59:57 +11:00
Sam	6380ebd829	FEATURE: allow personas to provide command options (#331 ) Personas now support providing options for commands. This PR introduces a single option "base_query" for the SearchCommand. When supplied all searches the persona will perform will also include the pre-supplied filter. This can allow personas to search a subset of the forum (such as documentation) This system is extensible we can add options to any command trivially.	2023-12-08 08:42:56 +11:00
Rafael dos Santos Silva	381b0d74ca	FIX: Handle truncation in HyDE search (#342 )	2023-12-07 10:36:56 -03:00
Roman Rizzi	450ec915d8	FIX: Make FoldContent strategy more resilient when using models with low token count. (#341 ) We'll recursively summarize the content into smaller chunks until we are sure we can concatenate them without going over the token limit.	2023-12-06 19:00:24 -03:00
Rafael dos Santos Silva	c8352f21ce	FIX: Fallback to whole LLM response when XML fail (#340 )	2023-12-06 18:58:26 -03:00
Rafael dos Santos Silva	252efdf142	FIX: Don't echo prompt back on HF/TGI (#338 ) * FIX: Don't echo prompt back on HF/TGI * teeeeests	2023-12-06 16:06:26 -03:00
Rafael dos Santos Silva	d8267d8da0	FIX: Many fixes for huggingface and llama2 inference (#335 )	2023-12-06 11:22:42 -03:00
Martin Brennan	24370a9ca6	Revert "FIX: Use Guardian.basic_user instead of new (anon) (#332 )" (#337 ) This reverts commit `a3a1285dc5`. c.f. https://github.com/discourse/discourse/pull/24742	2023-12-06 16:26:43 +10:00
Martin Brennan	a3a1285dc5	FIX: Use Guardian.basic_user instead of new (anon) (#332 ) c.f. de983796e1b66aa2ab039a4fb6e32cec8a65a098 There will soon be additional login_required checks for Guardian, and the intent of many checks by automated systems is better fulfilled by using BasicUser, which simulates a logged in TL0 forum user, rather than an anon user.	2023-12-06 12:01:41 +10:00
Rafael dos Santos Silva	71c5077228	FEATURE: User sentiment on profile summary page (#329 ) * FEATURE: User sentiment on profile summary page This introduces a new user stat in a user profile summary page. It will show either neutral/positive/negative according to the dominant sentiment in the user last interactions. The user-stat widget is only rendered for staff. Co-authored-by: Keegan George <kgeorge13@gmail.com>	2023-12-04 18:17:43 -03:00
Roman Rizzi	3bc010b686	FIX: call the right method to summarize with truncation (#328 )	2023-12-01 10:17:24 -03:00
Sam	a0b9fb9721	FIX: explicitly load embedding strategies (#325 ) If not, sometimes during tests these constants may not be loaded leading to flaky tests	2023-11-29 16:36:56 +11:00
Sam	6ddc17fd61	DEV: port directory structure to Zeitwerk (#319 ) Previous to this change we relied on explicit loading for a files in Discourse AI. This had a few downsides: - Busywork whenever you add a file (an extra require relative) - We were not keeping to conventions internally ... some places were OpenAI others are OpenAi - Autoloader did not work which lead to lots of full application broken reloads when developing. This moves all of DiscourseAI into a Zeitwerk compatible structure. It also leaves some minimal amount of manual loading (automation - which is loading into an existing namespace that may or may not be there) To avoid needing /lib/discourse_ai/... we mount a namespace thus we are able to keep /lib pointed at ::DiscourseAi Various files were renamed to get around zeitwerk rules and minimize usage of custom inflections Though we can get custom inflections to work it is not worth it, will require a Discourse core patch which means we create a hard dependency.	2023-11-29 15:17:46 +11:00
Rafael dos Santos Silva	fd0fb58eca	FEATURE: HuggingFace Text Embeddings Inference compatibility (#323 ) * FEATURE: HuggingFace Text Embeddings Inference compatibility * lint	2023-11-28 17:05:26 -03:00
Roman Rizzi	f26adf2cf6	FIX: Use XML tags in generate_titles prompt. (#322 ) We must ensure we can isolate titles, and the models sometimes ignore the example we give them. Additionally, anons can generate HyDE posts, so we need to check if user is nil when attempting to log requests.	2023-11-28 12:52:22 -03:00
Rafael dos Santos Silva	11e531b099	FEATURE: Backfill task for sentiment module (#316 ) * FEATURE: Backfill task for sentiment module * fix join clause	2023-11-28 12:28:36 -03:00
Roman Rizzi	2e7c5f047d	DEV: Don't attempt to update log if completion request fails. (#321 ) We already log the request failure when we raise the exception.	2023-11-28 11:15:12 -03:00
Roman Rizzi	775610b1c2	FIX: Chat titler was still using the old code after LLM migration (#314 )	2023-11-27 13:03:24 -03:00
Roman Rizzi	54a8dd9556	REFACTOR: Use LLM abstraction in the AI Helper. (#312 ) It also removes the need for multiple versions of our seeded prompts per model, further simplifying the code.	2023-11-27 09:33:31 -03:00
Sam	5a4598a7b4	FEATURE: Azure OpenAI support for DALLE 3 (#313 ) FEATURE: Azure OpenAI support for DALLE 3 Previous to this there was no way to add an inference endpoint for DALLE on Azure cause it requires custom URLs Also: - On save, when editing a persona it would revert priority and enabled - More forgiving parsing in command framework for array function calls - By default generate HD images - they tend to be a bit better - Improve DALL*E prompt which was getting very annoying and always echoing what it is about to do - Add a bit of a sleep between retries on image generation - Fix error handling in image_command	2023-11-27 13:01:05 +11:00
Sam	dff9f33a97	FEATURE: DALL-E-3 persona for image generation (#311 ) * FIX: no selected persona should pick first prioritized one Previously we were looking at `.personaId` but there is only an id attribute so it failed * FEATURE: new DALL-E-3 persona This persona generates images using DALL-E-3 API and is enabled by default Keep in mind that we are still waiting on seeds/gen_id so we can not retain style consistently between turns. This will change as soon as a new Open AI API provides the missing parameters Co-authored-by: Martin Brennan <martin@discourse.org>	2023-11-24 18:08:08 +11:00
Sam	6282b6d21f	FIX: implement tools framework for Anthropic (#307 ) Previous to this changeset we used a custom system for tools/command support for Anthropic. We defined commands by using !command as a signal to execute it Following Anthropic Claude 2.1, there is an official supported syntax (beta) for tools execution. eg: ``` + <function_calls> + <invoke> + <tool_name>image</tool_name> + <parameters> + <prompts> + [ + "an oil painting", + "a cute fluffy orange", + "3 apple's", + "a cat" + ] + </prompts> + </parameters> + </invoke> + </function_calls> ``` This implements the spec per Anthropic, it should be stable enough to also work on other LLMs. Keep in mind that OpenAI is not impacted here at all, as it has its own custom system for function calls. Additionally: - Fixes the title system prompt so it works with latest Anthropic - Uses new spec for "system" messages by Anthropic - Tweak forum helper persona to guide Anthropic a tiny be better Overall results are pretty awesome and Anthropic Claude performs really well now on Discourse	2023-11-24 06:39:56 +11:00
Roman Rizzi	419c43592a	FIX: Make summaries more cohesive by tweaking prompt. (#310 ) Other changes: - Don't use Bedrock for non claude models if credentials are set. - Remove extra sentence from HyDE prompt.	2023-11-23 16:33:37 -03:00
Roman Rizzi	02efca162e	FIX: Bedrock uses slightly different model names * Revert "FIX: We don't need to prepend anthropic. to bedrock models (#308)" This reverts commit `8a01751991`. * FIX: Bedrock uses slightly different model names	2023-11-23 15:49:24 -03:00
Roman Rizzi	8a01751991	FIX: We don't need to prepend anthropic. to bedrock models (#308 )	2023-11-23 14:39:21 -03:00
Roman Rizzi	3064d4c288	REFACTOR: Summarization and HyDE now use an LLM abstraction. (#297 ) * DEV: One LLM abstraction to rule them all * REFACTOR: HyDE search uses new LLM abstraction * REFACTOR: Summarization uses the LLM abstraction * Updated documentation and made small fixes. Remove Bedrock claude-2 restriction	2023-11-23 12:58:54 -03:00
Keegan George	787cd1bf17	FIX: Error 500 from search with only filters (#304 )	2023-11-22 11:05:56 -08:00
Roman Rizzi	e0691e70e8	DEV: Updates to the summarization strategy API (#301 ) Introduced by discourse/discourse#24489 In the future, this change will let us log who requested the summary in the `AiApiAuditLog`.:	2023-11-21 13:27:35 -03:00
Sam	5b5edb22c6	FEATURE: UI to update ai personas on admin page (#290 ) Introduces a UI to manage customizable personas (admin only feature) Part of the change was some extensive internal refactoring: - AIBot now has a persona set in the constructor, once set it never changes - Command now takes in bot as a constructor param, so it has the correct persona and is not generating AIBot objects on the fly - Added a .prettierignore file, due to the way ALE is configured in nvim it is a pre-req for prettier to work - Adds a bunch of validations on the AIPersona model, system personas (artist/creative etc...) are all seeded. We now ensure - name uniqueness, and only allow certain properties to be touched for system personas. - (JS note) the client side design takes advantage of nested routes, the parent route for personas gets all the personas via this.store.findAll("ai-persona") then child routes simply reach into this model to find a particular persona. - (JS note) data is sideloaded into the ai-persona model the meta property supplied from the controller, resultSetMeta - This removes ai_bot_enabled_personas and ai_bot_enabled_chat_commands, both should be controlled from the UI on a per persona basis - Fixes a long standing bug in token accounting ... we were doing to_json.length instead of to_json.to_s.length - Amended it so {commands} are always inserted at the end unconditionally, no need to add it to the template of the system message as it just confuses things - Adds a concept of required_commands to stock personas, these are commands that must be configured for this stock persona to show up. - Refactored tests so we stop requiring inference_stubs, it was very confusing to need it, added to plugin.rb for now which at least is clearer - Migrates the persona selector to gjs --------- Co-authored-by: Joffrey JAFFEUX <j.jaffeux@gmail.com> Co-authored-by: Martin Brennan <martin@discourse.org>	2023-11-21 16:56:43 +11:00
Roman Rizzi	244a338558	UX: Graph surprise next to joy in post_emotion chart (#291 )	2023-11-10 14:41:50 -03:00
Sam	a4f419f54f	FEATURE: basic infrastructure for custom personas (#288 ) - New AiPersona model which can store custom personas - Persona are restricted via group security - They can contain custom system messages - They can support a list of commands optionally To avoid expensive DB calls in the serializer a Multisite friendly Hash was introduced (which can be expired on transaction commit)	2023-11-10 11:39:49 +11:00
Roman Rizzi	d0198c5c5b	FIX: Changes to the sentiment reports. (#289 ) This PR aims to clarify sentiment reports by replacing averages with a count of posts that have one of their values above a threshold (60), meaning we have some level of confidence they are, in fact, positive or negative. Same thing happen with post emotions, with the difference that a post can have multiple values above it (30). Additionally, we dropped the "Neutral" axis. We also reworded the tooltip next to each report title, and added an early return to signal we have no data available instead of displaying an empty chart.	2023-11-09 17:23:25 -03:00
Roman Rizzi	458e66aef9	FIX: Filter classification type using the correct column (#286 )	2023-11-08 14:58:35 -03:00
Roman Rizzi	231cf91cc2	FIX: Don't divide by zero if there is no emotion data for TL group (#285 )	2023-11-08 13:05:36 -03:00
Roman Rizzi	b172ef11c4	FEATURE: Expose sentiment classifications via the admin dashboard. (#284 ) This PR adds new reports for displaying information about post sentiments grouped by date and emotions group by TL. Depends on discourse/discourse#24274	2023-11-08 10:50:37 -03:00
David Taylor	0902f74af5	DEV: Update linting configs (#280 )	2023-11-03 11:30:09 +00:00
Sam	fc65404896	FEATURE: support topic_id and post_id logging in ai audit log (#274 ) This makes it easier to track who is responsible for a completion in logs Note: ai helper and summarization are not yet implemented	2023-11-01 08:41:31 +11:00
Sam	0b62c0fa02	FIX: keep parity of shape for image command (#275 ) Function calling will start hallucinating if you reshape results. Previously we were morphing from: `{ prompts: ["prompt 1", "prompt 2"] }` to `{ prompts: { prompt: "prompt 1", seed: 222}, { ... ` This meant that over a few call sequences function_call starts hallucinating an incorrect shape. This change grounds us even on GPT-3.5	2023-10-31 19:12:25 +11:00
Rafael dos Santos Silva	6f708726e4	PERF: Better chat thread content format for LLM (#273 )	2023-10-30 19:57:46 -03:00
Rafael dos Santos Silva	cbfd8507b1	FIX: Update bedrock endpoint (#272 ) * FIX: Update bedrock endpoint AWS updated their endpoints per https://docs.aws.amazon.com/general/latest/gr/bedrock.html * lint	2023-10-30 19:27:50 -03:00
Rafael dos Santos Silva	0c9e18799c	FIX: unexpected return in aihelper entry_point (#271 )	2023-10-30 13:32:56 -03:00
Rafael dos Santos Silva	3c55ea8fc0	FEATURE: Automatic Chat Thread titles (#269 ) * FEATURE: Automatic Chat Thread titles * do not gen title for empty threads * make it default disabled for now	2023-10-30 11:56:33 -03:00
Sam	b06380d9fa	FIX: avoid semicolons at the end of queries for SQL Helper (#268 ) This makes it easier to cut and paste snippets it is producing Also fine tune the prompt in an attempt to hone gpt 3.5 which is very finicky	2023-10-27 16:21:09 +11:00
Sam	6add06af8f	FEATURE: Make artist more creative (#266 ) This allows for 2 big features: 1. Artist can ship up to 4 prompts for image generation 2. Artist can regenerate images cause it is aware of seed This allows for iteration on images maintaining visual style	2023-10-27 14:48:12 +11:00
Rafael dos Santos Silva	818b20fb6f	FEATURE: Make embeddings turn-key (#261 ) To ease the administrative burden of enabling the embeddings model, this change introduces automatic backfill when the setting is enabled. It also moves the topic visit embedding creation to a lower priority queue in sidekiq and adds an option to skip embedding computation and persistence when we match on the digest.	2023-10-26 12:07:37 -03:00
Sam	426e348c8a	FIX: make stable diffusion multi site friendly (#265 ) Previous to this change image generation did not work on multisite There was a background thread generating the images and it was getting site settings from the default site in the cluster This also removes referer header which is not needed	2023-10-25 11:04:16 +11:00
Rafael dos Santos Silva	0e5764617a	FEATURE: AI helper on posts (#244 ) Adds an AI Helper function when selecting text while viewing a topic. --------- Co-authored-by: Keegan George <kgeorge13@gmail.com> Co-authored-by: Roman Rizzi <roman@discourse.org>	2023-10-23 11:41:36 -03:00
Sam	1500308437	FEATURE: defer creation of bot users (#258 ) Also fixes it so users without bot in header can send it messages. Previous to this change we would seed all bots with database seeds. This lead to lots of confusion for people who do not enable ai bot. Instead: 1. We do not seed any bots until user enables the ai_bot_enabled setting 2. If it is disabled we will a. If no messages were created by bot - delete it b. Otherwise we will deactivate account	2023-10-23 17:00:58 +11:00
Ty Correll	87c591bbc2	UX: unify ai representing icon (#257 ) This PR addresses the effort to use one icon representing discourse-ai. Removed discourse-sparkles from discourse-ai, now included in core ``vendor/assets/svg-icons/discourse-additional.svg``	2023-10-19 17:31:56 -05:00
Sam	f65e50bd9e	FIX: allow for blank fields in Google results (#255 ) Under certain cases, for example: ``` there is this japanese band called kirimi, tell me more about them, try searching 3 times and at least 2 times in japanese before answering. ``` Results come back with blank snippets. This adds protection so this is allowed and code does not simply blow up.	2023-10-19 14:44:59 +11:00
Roman Rizzi	38f383a1e0	FIX: Allowlist topic custom field used by AI Bot (#250 )	2023-10-11 19:14:19 -03:00
Roman Rizzi	919a87f8d7	FIX: Include OP when building title suggestion prompt. (#248 ) When the OP is the only post in the topic, we'll send a prompt without user content to the LLM, and the suggested title will make no sense.	2023-10-10 14:08:08 +11:00

1 2 3 4 5 ...

397 Commits