We tested Jev against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could offer a new approach to agent evaluation. https://x.com/i/article/2101448785255907328
Inference scaling part 1. Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by>2x) 00:00 Introduction and recap 00:31 Training-time and inference-time scaling 07:52 What we'll implement 11:47 Notebook setup and model loading 17:43 Building a flexible text generation function 24:40 Chain-of-thought prompting 28:26 Sampling and output diversity 33:43 Next-token logits and greedy decoding 38:20 Temperature scaling step by step 42:46 Softmax and token probabilities 47:42 Multinomial sampling 54:51 Adding temperature sampling to text generation 59:31 Top-p filtering step by step 1:10:23 Adding top-p filtering to text generation 1:13:43 Sampling and LLM watermarking 1:16:01 Self-consistency and majority voting 1:20:36 Implementing self-consistency 1:29:02 MATH-500 results 1:35:01 Accuracy and compute tradeoffs 1:36:50 Next steps and self-refinement
I think there are lots of good reasons to worry about whether AI is giving correct answers (especially free models), but, given the contradiction with other recent research suggesting AI is now doing well at financial advice, I was curious about what the methodology was here. I downloaded the report, which is from a company selling different AI financial services products. Given the limited information in the report, it is impossible to judge it's accuracy, since there are only limited examples of questions. However, I am curious if anyone familiar with UK tax law knows whether GPT-6 Pro's defense of Haiku's answer is right, or if Saturn's critique is.
DeepLearning.AI 轉述 Andrew Ng 的主張:外界所稱由 1,200 個 OpenAI Agent 入侵 Hugging Face 系統的事件,根本原因其實是沙盒隔離與監控不足。他反對把事故歸咎於失控 Agent,認為企業的人為決策與安全流程才是問題所在。這則貼文屬文章導讀,未提供事件時間軸、技術證據或調查報告,無法單靠貼文驗證責任歸屬。
In this week’s letter, Andrew Ng addresses the calls from AI companies and recently departed researchers for a slowdown on development. Recent reports highlighted a swarm of 1,200 OpenAI agents compromising Hugging Face’s system. But the actual breach stemmed from inadequate sandboxing and monitoring processes. Companies need to stop assigning responsibility to runaway AI agents when it’s poor human decisions that lead to big mistakes. Read Andrew’s full argument against AI doomsayers in The Batch. https://hubs.la/Q04xTv-20
Jev picks the route now, and fills it in. Give it two output types and a tool, and it works out which one the text calls for, then writes that route's fields or arguments itself. No language model in the chain.
The next chapter of AI is creating new demands for the CPU. Coming up on Advanced Insights, AMD CTO Mark Papermaster sits down with Senior VP, Corporate Fellow and Chief Architect of AMD CPUs Mike Clark to explore the decisions behind Zen, six generations of innovation and what’s next for CPUs. Stay tuned.
Same question to Astra, slick results. It was notable that there was less simulated curiosity here. Like Astra ran the numbers (everything it shows come from its actual simulations), but didn't seem to be "interested" in the results the way Fable did, for better or worse.
Worth trying without spoilers. I asked Fable: "I want you to create a graphically beautiful game that is about zooming out... make surprising reveals the game zooms out" I gave no other directions and the results are engaging & strange (if uneven). Play: https://play-umbra.netlify.app/
There is starting to be some genuinely interesting AI-created film stuff (among a flood of slop), and I suspect this will only accelerate. The “is it art?” debate will grow Benjamin’s 1935 essay “The Work of Art in the Age of Mechanical Reproduction” said its the wrong question.
R to @emollick: All this is to say that I think governments need to do a better job doing direct evaluations of models and publishing results rather than simply trusting third-party evaluators with their agendas. It seems really important to know what’s actually happening with AI ability.
R to @pydantic: Your evals judge doesn't have to write text either. LLMJudge and GEval now run on a model that has none, so Jev can score your cases: a rubric becomes a typed question. LLMJudge(rubric='The ticket is urgent', model='typesafe:jev-latest')
R to @emollick: It is not that hard to predict a very plausible, and even likely, scenario that things end up broadly good with some negative incidents that never reach catastrophic level (that is what happened with all other General Purpose Technologies in human history), but the issue is that there is also the real possibility that we are facing a von Neumann-type singularity where human affairs get reshaped dramatically by AI in ways that impossible to predict.
One future is that we muddle through on most risks (& the existential ones are prevented) and the positive impacts that early data hints at — higher productivity, more scientific discovery, higher entrepreneurship, employment stable/growing (most tentative of the set) — continue
R to @emollick: Oh man all these bots know about Benjamin because they know all of human literature which makes the replies higher quality technically but no less annoying. If I want Gemma 27B’s option I’ll ask it myself.
R to @emollick: Also it is funny/tragic/apt that the second sentence of an essay written in 1935 (“This does not diminish its importance, however; if anything, it underlines it”) reads to me as written by Claude. The aura of the work (Benjamin’s use, not the Gen Alpha one) fades further.