会议开始前,我先说一下这个会议的背景。First of all, no actual BBQ today. It just a skill share. 这是我在AIE上看到的一个前vercel员工的workshop,他通过walk through他自己开发的工作流,我发现其中提到的grill me很有意思,所以就分享给了Shahab,然后Shahab觉得这个很好,就让我定了这个会议。 这个skill背后的原理是他发现,也是绝大多数用过AI的人都有遇到的问题,就是LLM Model是有smart zone和dump zone的,当你持续的问AI问题之后,在某个时间节点,你会发现LLM变得很dumb。基于这个发现的话,他生成了这个grill me的skill,你们在网上可能会看到不同的版本,但是差别不大。它主要的目的就是在让AI开始执行之前,想达到一个share understanding,只有这个时候,才让AI开始干活。 另外,在每个问题上,你不用全部自己回答,AI都会根据自己的理解,给不同direction的选项,如果这些选项你都不同意的话,你可以给出自己的要求,这样作的好处是,其实每一个方向,都是一个潜在的response的版本,与其在等AI回复完,你在花时间检查,不如提前确认好,这样不仅仅节省了token,更重要的是节省了时间。尤其是复杂的问题,一般会问二三十个问题,每个问题都有不同的方向,这个组合就会有非常多潜在的response的版本,如果从response上去调整的话,在一定的时间点,模型会开始变得很傻。 基于这个灵感,我把最近要作的一个项目拿来试了一下。这周我有个Looker的knowledge share,刚开始有只有非常初步的想法,然后通过这个skill,AI开始问我一些列的问题,干开始的是scope,这个比较正常,但是到第二轮的时候,它问我关于这个knowledge share的strong why是什么,这是个非常好的问题,如果我没有提前跟他沟通的话,这个版本可能会走向介绍Looker在作dashboard上怎么怎么好,所以当我回复之后,它回了一个big reframe。后面的问题我就不一一过了,它大概问了18个问题左右就感觉reach a share understanding了。然后我就让它帮我生成PPT,结果生产的PPT就我完全是我想要的,我一个字都不需要改的那种。 我能想想,如果是常规的方式,我问个问题,AI给一版本,我review,AI在生成,然后可能最终我也会得到一个版本,但是一会会话更多的时候在这个上面,另外,甚至有可能在某个时间点hit the dumb zone,哪个时候我可能就要带着这个draft重新开一个session。
The workflow should operate as a closed evaluation flywheel: One terminology adjustment: experimentation belongs mainly in the offline stage. The CI/CD stage should primarily apply repeatable evaluation suites and release gates rather than conduct uncontrolled experiments. 1. Business Goal Objective Define the business process the AI system should improve—not merely the model capability being demonstrated. […]
Part I: Key Evaluation Insights from the AI Engineer YouTube Channel The AI Engineer community has published a wide range of talks, workshops and case studies on LLM evaluation. Although the speakers use different tools and terminology, their recommendations consistently converge on one central idea: Effective evaluations should connect real business failures, reproducible test cases, […]
一、AI Engineer 频道的 Evals 核心内容地图 1. Domain-specific Evals:如何从业务错误开始 最重要的两场是: 这些视频共同强调: 不要先问“用什么指标”,而要先问“我们的系统到底会在哪些真实业务场景中失败”。 Hamel Husain 的方法尤其重要:找到真正懂业务的 Domain Expert,让其对真实输出做 Pass/Fail 判断并写 critique,再从这些人工判断中逐步构建自动化 LLM Judge。(youtube.com) Hamel 与 Shreya 后续把这个过程总结为: 2. 从 Playground 到 Production 主要视频包括: 这里的核心是把 Evals 从一次性 Notebook 变成持续运行的工程系统: Braintrust 的 workshop 明确覆盖 offline eval、online eval、production logging、用户反馈和人工审核;可扩展 Eval pipeline 则强调把数据、实验、评分和生产 traces 连接起来。(youtube.com) 3. Product Evals:PM 和业务人员如何参与 主要视频包括: 这类视频强调,Evals […]
下面是一份可直接用于公司内部分享的英文 Knowledge Share。核心设计是: 每一级 AI 项目都继承前一级的 Evals,同时增加与新复杂度对应的评估层。项目越复杂,不是替换旧指标,而是在共同的 Eval Core 上增加新的 Eval Packs。 From Simple AI Workflows to Multi-Agent Systems A Layered Evaluation Framework for Enterprise AI Projects Suggested duration: 25–30 minutesTarget audience: Product managers, AI engineers, data scientists, platform engineers, domain experts, security teams, and business owners 1. Why Do We Need Different Eval Strategies? […]
yield conn is used in Python’s context manager pattern with the @contextmanager decorator. It’s a special way to manage resources (like database connections) automatically. How it works in your code: In database.py: In main.py: Execution Flow: Why use yield? ✅ Benefits: Feature With yield Without yield Auto-close connection ✅ Yes ❌ Manual Exception safety ✅ […]
Django Overview What is Django? Django is a high-level Python web framework that enables rapid development of secure and maintainable websites. It follows the MTV (Model-Template-View) pattern (similar to MVC). Key Features: Philosophy: “The web framework for perfectionists with deadlines” Django vs Flask Feature Django Flask Type Full-stack framework Micro-framework Philosophy “Batteries included” Minimalist, flexible […]
Create virtual environment Start API app