discourse-ai/lib/completions/endpoints/fake.rb

# frozen_string_literal: true

module DiscourseAi
  module Completions
    module Endpoints
      class Fake < Base
        STOCK_CONTENT = <<~TEXT
      # Discourse Markdown Styles Showcase

      Welcome to the **Discourse Markdown Styles Showcase**! This _post_ is designed to demonstrate a wide range of Markdown capabilities available in Discourse.

      ## Lists and Emphasis

      - **Bold Text**: To emphasize a point, you can use bold text.
      - _Italic Text_: To subtly highlight text, italics are perfect.
      - ~~Strikethrough~~: Sometimes, marking text as obsolete requires a strikethrough.

      > **Note**: Combining these _styles_ can **_really_** make your text stand out!

      1. First item
      2. Second item
          * Nested bullet
          * Another nested bullet
      3. Third item

      ## Links and Images

      You can easily add [links](https://meta.discourse.org) to your posts. For adding images, use this syntax:

      ![Discourse Logo](https://meta.discourse.org/images/discourse-logo.svg)

      ## Code and Quotes

      Inline `code` is used for mentioning small code snippets like `let x = 10;`. For larger blocks of code, fenced code blocks are used:

      ```javascript
      function greet() {
          console.log("Hello, Discourse Community!");
      }
      greet();
      ```

      > Blockquotes can be very effective for highlighting user comments or important sections from cited sources. They stand out visually and offer great readability.

      ## Tables and Horizontal Rules

      Creating tables in Markdown is straightforward:

      | Header 1 | Header 2 | Header 3 |
      | ---------|:--------:| --------:|
      | Row 1, Col 1 | Centered | Right-aligned |
      | Row 2, Col 1 | **Bold** | _Italic_ |
      | Row 3, Col 1 | `Inline Code` | [Link](https://meta.discourse.org) |

      To separate content sections:

      ---

      ## Final Thoughts

      Congratulations, you've now seen a small sample of what Discourse's Markdown can do! For more intricate formatting, consider exploring the advanced styling options. Remember that the key to great formatting is not just the available tools, but also the **clarity** and **readability** it brings to your readers.
    TEXT

        def self.can_contact?(model_provider)
          model_provider == "fake"
        end

        def self.with_fake_content(content)
          @fake_content = content
          yield
        ensure
          @fake_content = nil
        end

        def self.fake_content=(content)
          @fake_content = content
        end

        def self.fake_content
          @fake_content || STOCK_CONTENT
        end

        def self.delays
          @delays ||= Array.new(10) { rand * 6 }
        end

        def self.delays=(delays)
          @delays = delays
        end

        def self.chunk_count
          @chunk_count ||= 10
        end

        def self.chunk_count=(chunk_count)
          @chunk_count = chunk_count
        end

        def self.last_call
          @last_call
        end

        def self.last_call=(params)
          @last_call = params
        end

        def self.previous_calls
          @previous_calls ||= []
        end

        def self.reset!
          @last_call = nil
          @fake_content = nil
          @delays = nil
          @chunk_count = nil
        end

        def perform_completion!(
          dialect,
          user,
          model_params = {},
          feature_name: nil,
          feature_context: nil
        )
          last_call = { dialect: dialect, user: user, model_params: model_params }
          self.class.last_call = last_call
          self.class.previous_calls << last_call
          # guard memory in test
          self.class.previous_calls.shift if self.class.previous_calls.length > 10

          content = self.class.fake_content

          content = content.shift if content.is_a?(Array)

          if block_given?
            if content.is_a?(DiscourseAi::Completions::ToolCall)
              yield(content, -> {})
            else
              split_indices = (1...content.length).to_a.sample(self.class.chunk_count - 1).sort
              indexes = [0, *split_indices, content.length]

              original_content = content
              content = +""

              cancel = false
              cancel_proc = -> { cancel = true }

              i = 0
              indexes
                .each_cons(2)
                .map { |start, finish| original_content[start...finish] }
                .each do |chunk|
                  break if cancel
                  if self.class.delays.present? &&
                       (delay = self.class.delays[i % self.class.delays.length])
                    sleep(delay)
                    i += 1
                  end
                  break if cancel

                  content << chunk
                  yield(chunk, cancel_proc)
                end
            end
          end

          content
        end
      end
    end
  end
end
FEATURE: smooth streaming of AI responses on the client (#413) This PR introduces 3 things: 1. Fake bot that can be used on local so you can test LLMs, to enable on dev use: SiteSetting.ai_bot_enabled_chat_bots = "fake" 2. More elegant smooth streaming of progress on LLM completion This leans on JavaScript to buffer and trickle llm results through. It also amends it so the progress dot is much more consistently rendered 3. It fixes the Claude dialect Claude needs newlines exactly at the right spot, amended so it is happy --------- Co-authored-by: Martin Brennan <martin@discourse.org> 2024-01-10 23:56:40 -05:00			`# frozen_string_literal: true`

			`module DiscourseAi`
			`module Completions`
			`module Endpoints`
			`class Fake < Base`
			`STOCK_CONTENT = <<~TEXT`
			`# Discourse Markdown Styles Showcase`

			`Welcome to the Discourse Markdown Styles Showcase! This _post_ is designed to demonstrate a wide range of Markdown capabilities available in Discourse.`

			`## Lists and Emphasis`

			`- Bold Text: To emphasize a point, you can use bold text.`
			`- _Italic Text_: To subtly highlight text, italics are perfect.`
			`- ~~Strikethrough~~: Sometimes, marking text as obsolete requires a strikethrough.`

			`> Note: Combining these _styles_ can _really_ make your text stand out!`

			`1. First item`
			`2. Second item`
			`* Nested bullet`
			`* Another nested bullet`
			`3. Third item`

			`## Links and Images`

			`You can easily add [links](https://meta.discourse.org) to your posts. For adding images, use this syntax:`

			`![Discourse Logo](https://meta.discourse.org/images/discourse-logo.svg)`

			`## Code and Quotes`

			Inline `code` is used for mentioning small code snippets like `let x = 10;`. For larger blocks of code, fenced code blocks are used:

			```javascript
			`function greet() {`
			`console.log("Hello, Discourse Community!");`
			`}`
			`greet();`
			```

			`> Blockquotes can be very effective for highlighting user comments or important sections from cited sources. They stand out visually and offer great readability.`

			`## Tables and Horizontal Rules`

			`Creating tables in Markdown is straightforward:`

			`\| Header 1 \| Header 2 \| Header 3 \|`
			`\| ---------\|:--------:\| --------:\|`
			`\| Row 1, Col 1 \| Centered \| Right-aligned \|`
			`\| Row 2, Col 1 \| Bold \| _Italic_ \|`
			\| Row 3, Col 1 \| `Inline Code` \| [Link](https://meta.discourse.org) \|

			`To separate content sections:`

			`---`

			`## Final Thoughts`

			`Congratulations, you've now seen a small sample of what Discourse's Markdown can do! For more intricate formatting, consider exploring the advanced styling options. Remember that the key to great formatting is not just the available tools, but also the clarity and readability it brings to your readers.`
			`TEXT`

DEV: Remove old code now that features rely on LlmModels. (#729) * DEV: Remove old code now that features rely on LlmModels. * Hide old settings and migrate persona llm overrides * Remove shadowing special URL + seeding code. Use srv:// prefix instead. 2024-07-30 12:44:57 -04:00			`def self.can_contact?(model_provider)`
			`model_provider == "fake"`
			`end`

FEATURE: Add Question Consolidator for robust Upload support in Personas (#596) This commit introduces a new feature for AI Personas called the "Question Consolidator LLM". The purpose of the Question Consolidator is to consolidate a user's latest question into a self-contained, context-rich question before querying the vector database for relevant fragments. This helps improve the quality and relevance of the retrieved fragments. Previous to this change we used the last 10 interactions, this is not ideal cause the RAG would "lock on" to an answer. EG: - User: how many cars are there in europe - Model: detailed answer about cars in europe including the term car and vehicle many times - User: Nice, what about trains are there in the US In the above example "trains" and "US" becomes very low signal given there are pages and pages talking about cars and europe. This mean retrieval is sub optimal. Instead, we pass the history to the "question consolidator", it would simply consolidate the question to "How many trains are there in the United States", which would make it fare easier for the vector db to find relevant content. The llm used for question consolidator can often be less powerful than the model you are talking to, we recommend using lighter weight and fast models cause the task is very simple. This is configurable from the persona ui. This PR also removes support for {uploads} placeholder, this is too complicated to get right and we want freedom to shift RAG implementation. Key changes: 1. Added a new `question_consolidator_llm` column to the `ai_personas` table to store the LLM model used for question consolidation. 2. Implemented the `QuestionConsolidator` module which handles the logic for consolidating the user's latest question. It extracts the relevant user and model messages from the conversation history, truncates them if needed to fit within the token limit, and generates a consolidated question prompt. 3. Updated the `Persona` class to use the Question Consolidator LLM (if configured) when crafting the RAG fragments prompt. It passes the conversation context to the consolidator to generate a self-contained question. 4. Added UI elements in the AI Persona editor to allow selecting the Question Consolidator LLM. Also made some UI tweaks to conditionally show/hide certain options based on persona configuration. 5. Wrote unit tests for the QuestionConsolidator module and updated existing persona tests to cover the new functionality. This feature enables AI Personas to better understand the context and intent behind a user's question by consolidating the conversation history into a single, focused question. This can lead to more relevant and accurate responses from the AI assistant. 2024-04-29 23:49:21 -04:00			`def self.with_fake_content(content)`
			`@fake_content = content`
			`yield`
			`ensure`
			`@fake_content = nil`
			`end`

FEATURE: new endpoint for directly accessing a persona (#876) The new `/admin/plugins/discourse-ai/ai-personas/stream-reply.json` was added. This endpoint streams data direct from a persona and can be used to access a persona from remote systems leaving a paper trail in PMs about the conversation that happened This endpoint is only accessible to admins. --------- Co-authored-by: Gabriel Grubba <70247653+Grubba27@users.noreply.github.com> Co-authored-by: Keegan George <kgeorge13@gmail.com> 2024-10-29 19:28:20 -04:00			`def self.fake_content=(content)`
			`@fake_content = content`
			`end`

FEATURE: smooth streaming of AI responses on the client (#413) This PR introduces 3 things: 1. Fake bot that can be used on local so you can test LLMs, to enable on dev use: SiteSetting.ai_bot_enabled_chat_bots = "fake" 2. More elegant smooth streaming of progress on LLM completion This leans on JavaScript to buffer and trickle llm results through. It also amends it so the progress dot is much more consistently rendered 3. It fixes the Claude dialect Claude needs newlines exactly at the right spot, amended so it is happy --------- Co-authored-by: Martin Brennan <martin@discourse.org> 2024-01-10 23:56:40 -05:00			`def self.fake_content`
			`@fake_content \|\| STOCK_CONTENT`
			`end`

			`def self.delays`
			`@delays \|\|= Array.new(10) { rand * 6 }`
			`end`

			`def self.delays=(delays)`
			`@delays = delays`
			`end`

			`def self.chunk_count`
			`@chunk_count \|\|= 10`
			`end`

			`def self.chunk_count=(chunk_count)`
			`@chunk_count = chunk_count`
			`end`

FEATURE: allow personas to supply top_p and temperature params (#459) * FEATURE: allow personas to supply top_p and temperature params Code assistance generally are more focused at a lower temperature This amends it so SQL Helper runs at 0.2 temperature vs the more common default across LLMs of 1.0. Reduced temperature leads to more focused, concise and predictable answers for the SQL Helper * fix tests * This is not perfect, but far better than what we do today Instead of fishing for 1. Draft sequence 2. Draft body We skip (2), this means the composer "only" needs 1 http request to open, we also want to eliminate (1) but it is a bit of a trickier core change, may figure out how to pull it off (defer it to first draft save) Value of bot drafts < value of opening bot conversations really fast 2024-02-02 15:09:34 -05:00			`def self.last_call`
			`@last_call`
			`end`

			`def self.last_call=(params)`
			`@last_call = params`
			`end`

FEATURE: support custom instructions for persona streaming (#890) This allows us to inject information into the system prompt which can help shape replies without repeating over and over in messages. 2024-11-04 15:43:26 -05:00			`def self.previous_calls`
			`@previous_calls \|\|= []`
			`end`

FEATURE: new endpoint for directly accessing a persona (#876) The new `/admin/plugins/discourse-ai/ai-personas/stream-reply.json` was added. This endpoint streams data direct from a persona and can be used to access a persona from remote systems leaving a paper trail in PMs about the conversation that happened This endpoint is only accessible to admins. --------- Co-authored-by: Gabriel Grubba <70247653+Grubba27@users.noreply.github.com> Co-authored-by: Keegan George <kgeorge13@gmail.com> 2024-10-29 19:28:20 -04:00			`def self.reset!`
			`@last_call = nil`
			`@fake_content = nil`
			`@delays = nil`
			`@chunk_count = nil`
			`end`

FEATURE: better logging for automation reports (#853) A new feature_context json column was added to ai_api_audit_logs This allows us to store rich json like context on any LLM request made. This new field now stores automation id and name. Additionally allows llm_triage to specify maximum number of tokens This means that you can limit the cost of llm triage by scanning only first N tokens of a post. 2024-10-23 01:49:56 -04:00			`def perform_completion!(`
			`dialect,`
			`user,`
			`model_params = {},`
			`feature_name: nil,`
			`feature_context: nil`
			`)`
FEATURE: support custom instructions for persona streaming (#890) This allows us to inject information into the system prompt which can help shape replies without repeating over and over in messages. 2024-11-04 15:43:26 -05:00			`last_call = { dialect: dialect, user: user, model_params: model_params }`
			`self.class.last_call = last_call`
			`self.class.previous_calls << last_call`
			`# guard memory in test`
			`self.class.previous_calls.shift if self.class.previous_calls.length > 10`
FEATURE: allow personas to supply top_p and temperature params (#459) * FEATURE: allow personas to supply top_p and temperature params Code assistance generally are more focused at a lower temperature This amends it so SQL Helper runs at 0.2 temperature vs the more common default across LLMs of 1.0. Reduced temperature leads to more focused, concise and predictable answers for the SQL Helper * fix tests * This is not perfect, but far better than what we do today Instead of fishing for 1. Draft sequence 2. Draft body We skip (2), this means the composer "only" needs 1 http request to open, we also want to eliminate (1) but it is a bit of a trickier core change, may figure out how to pull it off (defer it to first draft save) Value of bot drafts < value of opening bot conversations really fast 2024-02-02 15:09:34 -05:00
FEATURE: smooth streaming of AI responses on the client (#413) This PR introduces 3 things: 1. Fake bot that can be used on local so you can test LLMs, to enable on dev use: SiteSetting.ai_bot_enabled_chat_bots = "fake" 2. More elegant smooth streaming of progress on LLM completion This leans on JavaScript to buffer and trickle llm results through. It also amends it so the progress dot is much more consistently rendered 3. It fixes the Claude dialect Claude needs newlines exactly at the right spot, amended so it is happy --------- Co-authored-by: Martin Brennan <martin@discourse.org> 2024-01-10 23:56:40 -05:00			`content = self.class.fake_content`

FEATURE: new endpoint for directly accessing a persona (#876) The new `/admin/plugins/discourse-ai/ai-personas/stream-reply.json` was added. This endpoint streams data direct from a persona and can be used to access a persona from remote systems leaving a paper trail in PMs about the conversation that happened This endpoint is only accessible to admins. --------- Co-authored-by: Gabriel Grubba <70247653+Grubba27@users.noreply.github.com> Co-authored-by: Keegan George <kgeorge13@gmail.com> 2024-10-29 19:28:20 -04:00			`content = content.shift if content.is_a?(Array)`

FEATURE: smooth streaming of AI responses on the client (#413) This PR introduces 3 things: 1. Fake bot that can be used on local so you can test LLMs, to enable on dev use: SiteSetting.ai_bot_enabled_chat_bots = "fake" 2. More elegant smooth streaming of progress on LLM completion This leans on JavaScript to buffer and trickle llm results through. It also amends it so the progress dot is much more consistently rendered 3. It fixes the Claude dialect Claude needs newlines exactly at the right spot, amended so it is happy --------- Co-authored-by: Martin Brennan <martin@discourse.org> 2024-01-10 23:56:40 -05:00			`if block_given?`
FEATURE: improve tool support (#904) This re-implements tool support in DiscourseAi::Completions::Llm #generate Previously tool support was always returned via XML and it would be the responsibility of the caller to parse XML New implementation has the endpoints return ToolCall objects. Additionally this simplifies the Llm endpoint interface and gives it more clarity. Llms must implement decode, decode_chunk (for streaming) It is the implementers responsibility to figure out how to decode chunks, base no longer implements. To make this easy we ship a flexible json decoder which is easy to wire up. Also (new) Better debugging for PMs, we now have a next / previous button to see all the Llm messages associated with a PM Token accounting is fixed for vllm (we were not correctly counting tokens) 2024-11-11 16:14:30 -05:00			`if content.is_a?(DiscourseAi::Completions::ToolCall)`
			`yield(content, -> {})`
			`else`
			`split_indices = (1...content.length).to_a.sample(self.class.chunk_count - 1).sort`
			`indexes = [0, *split_indices, content.length]`

			`original_content = content`
			`content = +""`

			`cancel = false`
			`cancel_proc = -> { cancel = true }`

			`i = 0`
			`indexes`
			`.each_cons(2)`
			`.map { \|start, finish\| original_content[start...finish] }`
			`.each do \|chunk\|`
			`break if cancel`
			`if self.class.delays.present? &&`
			`(delay = self.class.delays[i % self.class.delays.length])`
			`sleep(delay)`
			`i += 1`
			`end`
			`break if cancel`

			`content << chunk`
			`yield(chunk, cancel_proc)`
FEATURE: smooth streaming of AI responses on the client (#413) This PR introduces 3 things: 1. Fake bot that can be used on local so you can test LLMs, to enable on dev use: SiteSetting.ai_bot_enabled_chat_bots = "fake" 2. More elegant smooth streaming of progress on LLM completion This leans on JavaScript to buffer and trickle llm results through. It also amends it so the progress dot is much more consistently rendered 3. It fixes the Claude dialect Claude needs newlines exactly at the right spot, amended so it is happy --------- Co-authored-by: Martin Brennan <martin@discourse.org> 2024-01-10 23:56:40 -05:00			`end`
FEATURE: improve tool support (#904) This re-implements tool support in DiscourseAi::Completions::Llm #generate Previously tool support was always returned via XML and it would be the responsibility of the caller to parse XML New implementation has the endpoints return ToolCall objects. Additionally this simplifies the Llm endpoint interface and gives it more clarity. Llms must implement decode, decode_chunk (for streaming) It is the implementers responsibility to figure out how to decode chunks, base no longer implements. To make this easy we ship a flexible json decoder which is easy to wire up. Also (new) Better debugging for PMs, we now have a next / previous button to see all the Llm messages associated with a PM Token accounting is fixed for vllm (we were not correctly counting tokens) 2024-11-11 16:14:30 -05:00			`end`
FEATURE: smooth streaming of AI responses on the client (#413) This PR introduces 3 things: 1. Fake bot that can be used on local so you can test LLMs, to enable on dev use: SiteSetting.ai_bot_enabled_chat_bots = "fake" 2. More elegant smooth streaming of progress on LLM completion This leans on JavaScript to buffer and trickle llm results through. It also amends it so the progress dot is much more consistently rendered 3. It fixes the Claude dialect Claude needs newlines exactly at the right spot, amended so it is happy --------- Co-authored-by: Martin Brennan <martin@discourse.org> 2024-01-10 23:56:40 -05:00			`end`

			`content`
			`end`
			`end`
			`end`
			`end`
			`end`