{"id":12910,"date":"2024-10-09T13:07:23","date_gmt":"2024-10-09T11:07:23","guid":{"rendered":"https:\/\/flaven.fr\/?p=12910"},"modified":"2026-09-16T10:09:58","modified_gmt":"2026-09-16T08:09:58","slug":"promptfoo-the-ultimate-tool-for-ensuring-llm-quality-and-reliability","status":"publish","type":"post","link":"https:\/\/flaven.fr\/2024\/10\/promptfoo-the-ultimate-tool-for-ensuring-llm-quality-and-reliability\/","title":{"rendered":"Promptfoo: The Ultimate Tool for Ensuring LLM Quality and Reliability"},"content":{"rendered":"<p>After exploring MLflow&#8217;s &#8220;using Prompt Engineering&#8221; feature. A feature that allowed me to reach a satisfactory &#8220;prompt+LLM&#8221; combination level and was the subhect of my previous post. <\/p>\n<p>Check &#8220;Enhance LLM Prompt Quality and Results with MLflow Integration&#8221; at <a href=\"https:\/\/wp.me\/p3Vuhl-3lS\" target=\"_blank\" rel=\"noopener\">https:\/\/wp.me\/p3Vuhl-3lS<\/a><\/p>\n<p><b>I was wondering: How to automatically test LLMs output to ensure that the quality\/relevance should be always present? So, you can ship LLM application online safely.<\/b><\/p>\n<p>This question is partly answered by &#8220;promptfoo&#8221; as its purpose is to &#8220;Test &#038; secure your LLM apps&#8221; <\/p>\n<p>Source : <a href=\"https:\/\/www.promptfoo.dev\/\" target=\"_blank\" rel=\"noopener\">https:\/\/www.promptfoo.dev\/<\/a><\/p>\n<p><b>For this post also, you can find all files and prompts, on my GitHub account. See <a href=\"https:\/\/github.com\/bflaven\/ia_usages\/tree\/main\/ia_testing_llm\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/bflaven\/ia_usages\/tree\/main\/ia_testing_llm<\/a><\/b><\/p>\n<p>The answer is simple: &#8220;promptfoo&#8221; has the ability in the same test, to simultaneously ensure the validity of the result produced by the LLM both in form and content.<\/p>\n<p><img width=\"600\" height=\"129\" src=\"https:\/\/flaven.fr\/wp-content\/uploads\/2024\/10\/promtfoo_cycle_600x129.png\" alt=\"Enhance LLM Prompt Quality and Results with MLflow Integration\" decoding=\"async\" fetchpriority=\"high\"><br \/>\n<i>Source: promptfoo<\/i><\/p>\n<p>So, what does &#8220;promptfoo&#8221; bring to the table? My testing scenario was the following:<\/p>\n<ul>\n<li>The first part of the test would be to ensure, for example, that the LLM outputs a valid JSON using a JSON model. Indeed, obtaining Structural Output is extremely useful for LLM-based applications since the process of parsing the results is much easier.<\/li>\n<li>The second part of the test would be to ensure, for example, that the generative AI produces content in the desired language and meets some validation criteria such as: the summary contains between 2 and 3 sentences, the number of desired keywords is five&#8230; etc.<\/li>\n<\/ul>\n<p>Obviously, these tests, at the LLM level, would be a useful complement to functional tests carried out with Cypress on the API, therefore on the application side.<\/p>\n<p>The combination of the two test suites would therefore offer a guarantee of quality on the results generated by the AI, within a CI\/CD.<\/p>\n<p>Technically, my POC is based on an LLM (Mistral) operated by Ollama. So I modified the `promptfooconfig.yaml` file to configure the `providers` so that everything is compatible with my development environment.<\/p>\n<h2>Good ressources on how-to use &#8220;promptfoo&#8221;<\/h2>\n<p>Here some good ressources that gave good insights on how to configure and use promptfoo:<\/p>\n<p>&#8220;How to Use Promptfoo for LLM Testing&#8221;, a good post as an introdution to promptfoo is this one: <a href=\"https:\/\/medium.com\/thedeephub\/how-to-use-promptfoo-for-llm-testing-13e96a9a9773\" target=\"_blank\" rel=\"noopener\">https:\/\/medium.com\/thedeephub\/how-to-use-promptfoo-for-llm-testing-13e96a9a9773<\/a><\/p>\n<p>A bunch of examples are provided by promptfoo so you can see the extent of possibilities offered by the framework. Check <a href=\"https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples<\/a><\/p>\n<p><b>For me, these 8 examples in particular have been a good source of inspiration.<\/b><\/p>\n<ol>\n<li>python-assert-external: <a href=\"https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/python-assert-external\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/python-assert-external<\/a><\/li>\n<li>json-output: <a href=\"https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/python-assert-external\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/python-assert-external<\/a><\/li>\n<li>prompts-per-model: <a href=\"https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/prompts-per-model\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/prompts-per-model<\/a><\/li>\n<li>summarization: <a href=\"https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/summarization\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/summarization<\/a><\/li>\n<li>simple-cli: <a href=\"https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/simple-cli\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/simple-cli<\/a><\/li>\n<li>simple-csv: <a href=\"https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/simple-csv\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/simple-csv<\/a><\/li>\n<li>simple-test: <a href=\"https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/simple-test\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/simple-test<\/a><\/li>\n<li>mistral-llama-comparison: <a href=\"https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/mistral-llama-comparison\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo\/tree\/main\/examples\/mistral-llama-comparison<\/a><\/li>\n<\/ol>\n<p><b>After reading this post and browse the examples, my &#8220;shopping&#8221; list was the following. I am operating the LLMs locally with the help of Ollama and the model I am using is Mistral.<\/b><\/p>\n<ol>\n<li>Add provider via id to be able to test the different models (mistral-openorca, mistral, openhermes, phi3, zephyr)<\/li>\n<li>Externalize prompts, externalize contents (articles)<\/li>\n<li>Do a test on the json model output from LLM<\/li>\n<li>Do a test using python to qualify the quality of the result from a relevance point of view.<\/li>\n<\/ol>\n<p><b>Below an extract on how to set the provider in the file promptfooconfig.yaml<\/b><\/p>\n<pre>\r\nproviders:\r\n  - id: openrouter:mistralai\/mistral-7b-instruct\r\n    config:\r\n      temperature: 0.5\r\n  - id: openrouter:mistralai\/mixtral-8x7b-instruct\r\n    config:\r\n      temperature: 0.5\r\n  - id: openrouter:meta-llama\/llama-3.1-8b-instruct\r\n    config:\r\n      temperature: 0.5\r\n<\/pre>\n<p><b>The walkthrough<\/b><\/p>\n<pre>\r\n\r\n#create the dir\r\nmkdir 001_promptfoo_running\r\n\r\n# path\r\ncd \/Users\/brunoflaven\/Documents\/01_work\/blog_articles\/ia_testing_llm\/001_promptfoo_running\/\r\n\r\n# install promptfoo\r\nnpm install -g promptfoo@latest\r\n\r\n# change providers in promptfooconfig.yaml\r\n- ollama:mistral:latest\r\n\r\n# eval\r\nnpx promptfoo eval\r\n\r\n# launch commands\r\nLOG_LEVEL=debug npx promptfoo eval\r\n\r\n# launch eval\r\nnpx promptfoo eval\r\n\r\n# view result\r\nnpx promptfoo view\r\n\r\n#uninstall\r\nnpm uninstall -g promptfoo\r\n# gain space disk\r\nnpm cache clean --force\r\n<\/pre>\n<p><b>Conclusion:<\/b><\/p>\n<p>Integrate Promptfoo into your development workflow or in your CI\/CD, this is probably the only way to enhance quality, and reliability of the LLM output.<\/p>\n<p>Here is below the praises sung by the promptfoo&#8217;s creators themselves :<\/p>\n<ul>\n<li><strong>Developer friendly<\/strong>: promptfoo is fast, with quality-of-life features like live reloads and caching.<\/li>\n<li><strong>Battle-tested<\/strong>: Originally built for LLM apps serving over 10 million users in production. Our tooling is flexible and can be adapted to many setups.<\/li>\n<li><strong>Simple, declarative test cases<\/strong>: Define evals without writing code or working with heavy notebooks.<\/li>\n<li><strong>Language agnostic<\/strong>: Use Python, Javascript, or any other language.<\/li>\n<li><strong>Share &amp; collaborate<\/strong>: Built-in share functionality &amp; web viewer for working with teammates.<\/li>\n<li><strong>Open-source<\/strong>: LLM evals are a commodity and should be served by 100% open-source projects with no strings attached.<\/li>\n<li><strong>Private<\/strong>: This software runs completely locally. The evals run on your machine and talk directly with the LLM.<\/li>\n<\/ul>\n<p>Source: <a href=\"https:\/\/www.promptfoo.dev\/docs\/intro\/\" target=\"_blank\" rel=\"noopener\">https:\/\/www.promptfoo.dev\/docs\/intro\/<\/a><\/p>\n<p><b>Extra: Using NotebookLM<\/b><\/p>\n<p>This post is also an expreiment to test NotebookLM. So, here is this regular blog post &#8220;Promptfoo: The Ultimate Tool for Ensuring LLM Quality and Reliability&#8221; converted into a podcast using NotebookLM.<\/p>\n<blockquote><p>NotebookLM gives you a personalized AI collaborator that helps you do your best thinking. After uploading your documents, NotebookLM becomes an instant expert in those sources so you can read, take notes, and collaborate with it to refine and organize your ideas.<\/p><\/blockquote>\n<p><iframe loading=\"lazy\" width=\"100%\" height=\"166\" scrolling=\"no\" frameborder=\"no\" allow=\"autoplay\" src=\"https:\/\/w.soundcloud.com\/player\/?url=https%3A\/\/api.soundcloud.com\/tracks\/1931591423&#038;color=%23ff5500&#038;auto_play=false&#038;hide_related=false&#038;show_comments=true&#038;show_user=true&#038;show_reposts=false&#038;show_teaser=true\"><\/iframe><\/p>\n<div style=\"font-size: 10px; color: #cccccc;line-break: anywhere;word-break: normal;overflow: hidden;white-space: nowrap;text-overflow: ellipsis; font-family: Interstate,Lucida Grande,Lucida Sans Unicode,Lucida Sans,Garuda,Verdana,Tahoma,sans-serif;font-weight: 100;\"><a href=\"https:\/\/soundcloud.com\/bruno-flaven\" title=\"Bruno Flaven\" target=\"_blank\" style=\"color: #cccccc; text-decoration: none;\" rel=\"noopener\">Bruno Flaven<\/a> \u00b7 <a href=\"https:\/\/soundcloud.com\/bruno-flaven\/ia_testing_llm_3\" title=\"ia_testing_llm_3\" target=\"_blank\" style=\"color: #cccccc; text-decoration: none;\" rel=\"noopener\">ia_testing_llm_3<\/a><\/div>\n<p>More on : <a href=\"https:\/\/support.google.com\/notebooklm\/#topic=14287611\" target=\"_blank\" rel=\"noopener\">https:\/\/support.google.com\/notebooklm\/#topic=14287611<\/a> <\/p>\n<h2>Videos to tackle this post<\/h2>\n<p>Promptfoo: The Ultimate Tool for Ensuring LLM Quality and Reliability (Part 1)<br \/>\n<iframe loading=\"lazy\" width=\"560\" height=\"315\" src=\"https:\/\/www.youtube.com\/embed\/hFh_DkN63KU\" title=\"YouTube video player\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture\" allowfullscreen><\/iframe><\/p>\n<p>Promptfoo: The Ultimate Tool for Ensuring LLM Quality and Reliability (Part 2)<br \/>\n<iframe loading=\"lazy\" width=\"560\" height=\"315\" src=\"https:\/\/www.youtube.com\/embed\/ZRuqwKowBWI\" title=\"YouTube video player\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture\" allowfullscreen><\/iframe><\/p>\n<p><H2>More infos<\/H2><\/p>\n<ul>\n<li>Testing Language Models (and Prompts) Like We Test Software | by Marco Tulio Ribeiro | Towards Data Science<br \/><a href=\"https:\/\/towardsdatascience.com\/testing-large-language-models-like-we-test-software-92745d28a359\" target=\"_blank\" rel=\"noopener\">https:\/\/towardsdatascience.com\/testing-large-language-models-like-we-test-software-92745d28a359<\/a><\/li>\n<li>GitHub &#8211; aws-samples\/llm-based-advanced-summarization<br \/><a href=\"https:\/\/github.com\/aws-samples\/llm-based-advanced-summarization\/tree\/main\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/aws-samples\/llm-based-advanced-summarization\/tree\/main<\/a><\/li>\n<li>\n        Prompt Engineering Testing Strategies with Python |<br \/>\n      Shiro<br \/>\n    <br \/><a href=\"https:\/\/openshiro.com\/articles\/prompt-engineering-testing-strategies-with-python\" target=\"_blank\" rel=\"noopener\">https:\/\/openshiro.com\/articles\/prompt-engineering-testing-strategies-with-python<\/a><\/li>\n<li>GitHub &#8211; duncantmiller\/llm_prompt_engineering: Prompt engineering testing strategies, using the OpenAI API.<br \/><a href=\"https:\/\/github.com\/duncantmiller\/llm_prompt_engineering\/tree\/main\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/duncantmiller\/llm_prompt_engineering\/tree\/main<\/a><\/li>\n<li>Generative AI Evaluation with Promptfoo: A Comprehensive Guide | by Yuki Nagae | Sep, 2024 | Medium<br \/><a href=\"https:\/\/medium.com\/@yukinagae\/generative-ai-evaluation-with-promptfoo-a-comprehensive-guide-e23ea95c1bb7\" target=\"_blank\" rel=\"noopener\">https:\/\/medium.com\/@yukinagae\/generative-ai-evaluation-with-promptfoo-a-comprehensive-guide-e23ea95c1bb7<\/a><\/li>\n<li>promptfoo\/examples\/assistant-cli at 0665ec88d58369ede2d0615c24c1f023b7fafa9b \u00b7 promptfoo\/promptfoo \u00b7 GitHub<br \/><a href=\"https:\/\/github.com\/promptfoo\/promptfoo\/tree\/0665ec88d58369ede2d0615c24c1f023b7fafa9b\/examples\/assistant-cli\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo\/tree\/0665ec88d58369ede2d0615c24c1f023b7fafa9b\/examples\/assistant-cli<\/a><\/li>\n<li>Getting started | promptfoo<br \/><a href=\"https:\/\/www.promptfoo.dev\/docs\/getting-started\/\" target=\"_blank\" rel=\"noopener\">https:\/\/www.promptfoo.dev\/docs\/getting-started\/<\/a><\/li>\n<li>Model-graded metrics | promptfoo<br \/><a href=\"https:\/\/www.promptfoo.dev\/docs\/configuration\/expected-outputs\/model-graded\/\" target=\"_blank\" rel=\"noopener\">https:\/\/www.promptfoo.dev\/docs\/configuration\/expected-outputs\/model-graded\/<\/a><\/li>\n<li>Solving the \u201cPunycode Module is Deprecated\u201d Issue in Node.js | by Asimabas | Aug, 2024 | Medium<br \/><a href=\"https:\/\/medium.com\/@asimabas96\/solving-the-punycode-module-is-deprecated-issue-in-node-js-93437637948a\" target=\"_blank\" rel=\"noopener\">https:\/\/medium.com\/@asimabas96\/solving-the-punycode-module-is-deprecated-issue-in-node-js-93437637948a<\/a><\/li>\n<li>Semantic Tagging: Create Meaningful Tags for your Text Data | by Gabriele Sgroi, PhD | Towards AI<br \/><a href=\"https:\/\/pub.towardsai.net\/semantic-tagging-create-meaningful-tags-for-your-text-data-dcf8d2f24960\" target=\"_blank\" rel=\"noopener\">https:\/\/pub.towardsai.net\/semantic-tagging-create-meaningful-tags-for-your-text-data-dcf8d2f24960<\/a><\/li>\n<li>How to use Large Language Models to tag your data: A complete tutorial | by Research Graph | Medium<br \/><a href=\"https:\/\/medium.com\/@researchgraph\/how-to-use-large-language-models-to-tag-your-data-a-complete-tutorial-4a3647ae0f05\" target=\"_blank\" rel=\"noopener\">https:\/\/medium.com\/@researchgraph\/how-to-use-large-language-models-to-tag-your-data-a-complete-tutorial-4a3647ae0f05<\/a><\/li>\n<li>Welcome To Instructor &#8211; Instructor<br \/><a href=\"https:\/\/jxnl.github.io\/instructor\/\" target=\"_blank\" rel=\"noopener\">https:\/\/jxnl.github.io\/instructor\/<\/a><\/li>\n<li>Ollama &#8211; Instructor<br \/><a href=\"https:\/\/jxnl.github.io\/instructor\/examples\/ollama\/#ollama\" target=\"_blank\" rel=\"noopener\">https:\/\/jxnl.github.io\/instructor\/examples\/ollama\/#ollama<\/a><\/li>\n<li>GitHub &#8211; guidance-ai\/guidance: A guidance language for controlling large language models.<br \/><a href=\"https:\/\/github.com\/guidance-ai\/guidance\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/guidance-ai\/guidance<\/a><\/li>\n<li>Evaluate AI\/LLM Performance with Effective Test Prompts<br \/><a href=\"https:\/\/writingmate.ai\/blog\/ai-llm-perfomance-testing\" target=\"_blank\" rel=\"noopener\">https:\/\/writingmate.ai\/blog\/ai-llm-perfomance-testing<\/a><\/li>\n<li>GitHub &#8211; promptfoo\/promptfoo: Test your prompts, agents, and RAGs. Red teaming, pentesting, and vulnerability scanning for LLMs. Compare performance of GPT, Claude, Gemini, Llama, and more. Simple declarative configs with command line and CI\/CD integration.<br \/><a href=\"https:\/\/github.com\/promptfoo\/promptfoo\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/promptfoo\/promptfoo<\/a><\/li>\n<li>LLM evaluation techniques for JSON outputs | promptfoo<br \/><a href=\"https:\/\/www.promptfoo.dev\/docs\/guides\/evaluate-json\/\" target=\"_blank\" rel=\"noopener\">https:\/\/www.promptfoo.dev\/docs\/guides\/evaluate-json\/<\/a><\/li>\n<li>Test Driven PROMPT Engineering: Using Promptfoo to COMPARE Prompts, LLMs, and Providers. &#8211; YouTube<br \/><a href=\"https:\/\/www.youtube.com\/watch?v=KhINc5XwhKs\" target=\"_blank\" rel=\"noopener\">https:\/\/www.youtube.com\/watch?v=KhINc5XwhKs<\/a><\/li>\n<li>GitHub &#8211; disler\/llm-prompt-testing-quick-start: LLM Prompt Testing Quick Start<br \/><a href=\"https:\/\/github.com\/disler\/llm-prompt-testing-quick-start\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/disler\/llm-prompt-testing-quick-start<\/a><\/li>\n<li>Phi vs Llama: Benchmark on your own data | promptfoo<br \/><a href=\"https:\/\/www.promptfoo.dev\/docs\/guides\/phi-vs-llama\/\" target=\"_blank\" rel=\"noopener\">https:\/\/www.promptfoo.dev\/docs\/guides\/phi-vs-llama\/<\/a><\/li>\n<li>Enhancing JSON Output with Large Language Models: A Comprehensive Guide | by Dina Berenbaum | Medium<br \/><a href=\"https:\/\/medium.com\/@dinber19\/enhancing-json-output-with-large-language-models-a-comprehensive-guide-f1935aa724fb\" target=\"_blank\" rel=\"noopener\">https:\/\/medium.com\/@dinber19\/enhancing-json-output-with-large-language-models-a-comprehensive-guide-f1935aa724fb<\/a><\/li>\n<li>How to Get Only JSON response from Any LLM Using LangChain | by Harshit Dubey | Medium<br \/><a href=\"https:\/\/medium.com\/@harshitdy\/how-to-get-only-json-response-from-any-llm-using-langchain-ed53bc2df50f\" target=\"_blank\" rel=\"noopener\">https:\/\/medium.com\/@harshitdy\/how-to-get-only-json-response-from-any-llm-using-langchain-ed53bc2df50f<\/a><\/li>\n<li>LLM Evaluation: Comparing Four Methods to Automatically Detect Errors | Label Studio<br \/><a href=\"https:\/\/labelstud.io\/blog\/llm-evaluation-comparing-four-methods-to-automatically-detect-errors\/?trk=public_post_comment-text\" target=\"_blank\" rel=\"noopener\">https:\/\/labelstud.io\/blog\/llm-evaluation-comparing-four-methods-to-automatically-detect-errors\/?trk=public_post_comment-text<\/a><\/li>\n<li>Your AI Product Needs Evals \u2013 Hamel&#8217;s Blog<br \/><a href=\"https:\/\/hamel.dev\/blog\/posts\/evals\/\" target=\"_blank\" rel=\"noopener\">https:\/\/hamel.dev\/blog\/posts\/evals\/<\/a><\/li>\n<li>GitHub &#8211; confident-ai\/deepeval: The LLM Evaluation Framework<br \/><a href=\"https:\/\/github.com\/confident-ai\/deepeval\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/confident-ai\/deepeval<\/a><\/li>\n<li>LLM Observability &#038; Application Tracing (open source) &#8211; Langfuse<br \/><a href=\"https:\/\/langfuse.com\/docs\/tracing\" target=\"_blank\" rel=\"noopener\">https:\/\/langfuse.com\/docs\/tracing<\/a><\/li>\n<li>Google Colab<br \/><a href=\"https:\/\/colab.research.google.com\/github\/langfuse\/langfuse-docs\/blob\/main\/cookbook\/integration_ollama.ipynb\" target=\"_blank\" rel=\"noopener\">https:\/\/colab.research.google.com\/github\/langfuse\/langfuse-docs\/blob\/main\/cookbook\/integration_ollama.ipynb<\/a><\/li>\n<li>llm-evaluation \u00b7 GitHub Topics \u00b7 GitHub<br \/><a href=\"https:\/\/github.com\/topics\/llm-evaluation\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/topics\/llm-evaluation<\/a><\/li>\n<li>prompt-testing \u00b7 GitHub Topics \u00b7 GitHub<br \/><a href=\"https:\/\/github.com\/topics\/prompt-testing\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/topics\/prompt-testing<\/a><\/li>\n<li>Collecting user feedback on ML in Streamlit<br \/><a href=\"https:\/\/blog.streamlit.io\/collecting-user-feedback-on-ml-in-streamlit\/\" target=\"_blank\" rel=\"noopener\">https:\/\/blog.streamlit.io\/collecting-user-feedback-on-ml-in-streamlit\/<\/a><\/li>\n<li>Trubrics<br \/><a href=\"https:\/\/www.trubrics.com\/#pricing\" target=\"_blank\" rel=\"noopener\">https:\/\/www.trubrics.com\/#pricing<\/a><\/li>\n<li>How to capture the feedback effectively &#8211; #3 by ferdy &#8211; Using Streamlit &#8211; Streamlit<br \/><a href=\"https:\/\/discuss.streamlit.io\/t\/how-to-capture-the-feedback-effectively\/60138\/3\" target=\"_blank\" rel=\"noopener\">https:\/\/discuss.streamlit.io\/t\/how-to-capture-the-feedback-effectively\/60138\/3<\/a><\/li>\n<li>Collect user feedback on AI models from your Streamlit app &#8211; YouTube<br \/><a href=\"https:\/\/www.youtube.com\/watch?v=2Qt54qGwIdQ\" target=\"_blank\" rel=\"noopener\">https:\/\/www.youtube.com\/watch?v=2Qt54qGwIdQ<\/a><\/li>\n<li>Getting Started | LMQL<br \/><a href=\"https:\/\/lmql.ai\/docs\/\" target=\"_blank\" rel=\"noopener\">https:\/\/lmql.ai\/docs\/<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>After exploring MLflow&#8217;s &#8220;using Prompt Engineering&#8221; feature. A feature that allowed me to reach a satisfactory &#8220;prompt+LLM&#8221; combination level and was the subhect of my&hellip; <\/p>\n<p class=\"text-center\"><a href=\"https:\/\/flaven.fr\/2024\/10\/promptfoo-the-ultimate-tool-for-ensuring-llm-quality-and-reliability\/\" class=\"more-link\">Continue reading &rarr; <span class=\"screen-reader-text\">Promptfoo: The Ultimate Tool for Ensuring LLM Quality and Reliability<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":12914,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"bf_ai_meta_description":"Promptfoo ensures LLM quality and reliability by testing outputs for validity and relevance, enabling safe online deployment.","bf_ai_og_title":"Promptfoo: LLM Quality Tool","footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[3438,3437,3444,3447,3448,3449,3450,3435],"tags":[3221,3223,3515,3493,2122,3219,3220,2886,3501,226,192,3076,141,3089,3224,3083,218,3517,3217,3227,3214,3225,3226,2316,3215,3216,3222,3218,2395],"class_list":["post-12910","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-machine-learning","category-business-case-studies","category-programming-databases","category-technology-trends","category-tools-productivity","category-tutorials-how-to","category-ux-product-design","category-web-development","tag-battle-tested","tag-ci-cd","tag-claude","tag-clip","tag-collaboration","tag-content-based-criteria","tag-developer-friendly","tag-examples","tag-google-colab","tag-javascript","tag-json","tag-llm","tag-local","tag-mistral","tag-mlflow","tag-ollama","tag-open-source","tag-openrouter","tag-output","tag-private","tag-promptfoo","tag-promptfooconfig-yaml","tag-providers","tag-python","tag-quality","tag-relevance","tag-reliability","tag-structural-criteria","tag-testing"],"jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/p3Vuhl-3me","jetpack_featured_media_url":"https:\/\/flaven.fr\/wp-content\/uploads\/2024\/10\/ia_testing_llm_b.png","_links":{"self":[{"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/posts\/12910","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/comments?post=12910"}],"version-history":[{"count":8,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/posts\/12910\/revisions"}],"predecessor-version":[{"id":12924,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/posts\/12910\/revisions\/12924"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/media\/12914"}],"wp:attachment":[{"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/media?parent=12910"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/categories?post=12910"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/tags?post=12910"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}