並び順

ブックマーク数

期間指定

  • から
  • まで

1 - 40 件 / 74件

新着順 人気順

computer problem solving process stepsの検索結果1 - 40 件 / 74件

  • GPT-5.1 Prompting Guide

    Introduction GPT-5.1, our newest flagship model, is designed to balance intelligence and speed for a variety of agentic and coding tasks, while also introducing a new none reasoning mode for low-latency interactions. Building on the strengths of GPT-5, GPT-5.1 is better calibrated to prompt difficulty, consuming far fewer tokens on easy inputs and more efficiently handling challenging ones. Along

      GPT-5.1 Prompting Guide
    • How we built our multi-agent research system

      Published Jun 13, 2025 Our Research feature uses multiple Claude agents to explore complex topics more effectively. We share the engineering challenges and the lessons we learned from building this system. Claude now has Research capabilities that allow it to search across the web, Google Workspace, and any integrations to accomplish complex tasks. The journey of this multi-agent system from proto

        How we built our multi-agent research system
      • Dario Amodei — The Adolescence of Technology

        There is a scene in the movie version of Carl Sagan’s book Contact where the main character, an astronomer who has detected the first radio signal from an alien civilization, is being considered for the role of humanity’s representative to meet the aliens. The international panel interviewing her asks, “If you could ask [the aliens] just one question, what would it be?” Her reply is: “I’d ask them

          Dario Amodei — The Adolescence of Technology
        • Things we learned about LLMs in 2024

          31st December 2024 A lot has happened in the world of Large Language Models over the course of 2024. Here’s a review of things we figured out about the field in the past twelve months, plus my attempt at identifying key themes and pivotal moments. This is a sequel to my review of 2023. In this article: The GPT-4 barrier was comprehensively broken Some of those GPT-4 models run on my laptop LLM pri

            Things we learned about LLMs in 2024
          • Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

            Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku Update (12/03/2024): We have revised the pricing for Claude 3.5 Haiku. The model is now priced at $0.80 MTok input / $4 MTok output. Today, we’re announcing an upgraded Claude 3.5 Sonnet, and a new model, Claude 3.5 Haiku. The upgraded Claude 3.5 Sonnet delivers across-the-board improvements over its predecessor, with particul

              Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
            • RFC 9562: Universally Unique IDentifiers (UUIDs)

               Internet Engineering Task Force (IETF) K. Davis Request for Comments: 9562 Cisco Systems Obsoletes: 4122 B. Peabody Category: Standards Track Uncloud ISSN: 2070-1721 P. Leach University of Washington May 2024 Universally Unique IDentifiers (UUIDs) Abstract This specification defines UUIDs (Universally Unique IDentifiers) -- also known as GUIDs (Globally Unique IDentifiers) -- and a Uniform Resou

                RFC 9562: Universally Unique IDentifiers (UUIDs)
              • Agents

                Intelligent agents are considered by many to be the ultimate goal of AI. The classic book by Stuart Russell and Peter Norvig, Artificial Intelligence: A Modern Approach (Prentice Hall, 1995), defines the field of AI research as “the study and design of rational agents.” The unprecedented capabilities of foundation models have opened the door to agentic applications that were previously unimaginabl

                  Agents
                • How we built our multi-agent research system

                  Published Jun 13, 2025 Our Research feature uses multiple Claude agents to explore complex topics more effectively. We share the engineering challenges and the lessons we learned from building this system. Claude now has Research capabilities that allow it to search across the web, Google Workspace, and any integrations to accomplish complex tasks. The journey of this multi-agent system from proto

                    How we built our multi-agent research system
                  • From Coder to Orchestrator: The future of software engineering with AI - Human Who Codes

                    The software engineering industry is undergoing a major AI-driven transition in how we work. The days when humans needed to write every line of code are already behind us as LLMs become more capable and reliable. The improvement in code output during 2025 alone has been astounding. I’ve personally watched LLMs struggle with certain problems, then a few months later, solve them completely and effic

                      From Coder to Orchestrator: The future of software engineering with AI - Human Who Codes
                    • How to Do Great Work

                      July 2023 If you collected lists of techniques for doing great work in a lot of different fields, what would the intersection look like? I decided to find out by making it. Partly my goal was to create a guide that could be used by someone working in any field. But I was also curious about the shape of the intersection. And one thing this exercise shows is that it does have a definite shape; it's

                      • claude-cycles.dvi

                        Claude’s Cycles Don Knuth, Stanford Computer Science Department (28 February 2026; revised 06 March 2026) Shock! Shock! I learned yesterday that an open problem I’d been working on for several weeks had just been solved by Claude Opus 4.6—Anthropic’s hybrid reasoning model that had been released three weeks earlier! It seems that I’ll have to revise my opinions about “generative AI” one of these d

                        • Patterns for Building LLM-based Systems & Products

                          Patterns for Building LLM-based Systems & Products [ llm engineering production 🔥 ] · 66 min read Discussions on HackerNews, Twitter, and LinkedIn “There is a large class of problems that are easy to imagine and build demos for, but extremely hard to make products out of. For example, self-driving: It’s easy to demo a car self-driving around a block, but making it into a product takes a decade.”

                            Patterns for Building LLM-based Systems & Products
                          • An Economy of AI Agents

                            An Economy of AI Agents Gillian K. Hadfield* Johns Hopkins Andrew Koh† MIT This version: September 3, 2025 Prepared for the NBER Handbook on the Economics of Transformative AI Abstract In the coming decade, artificially intelligent agents with the ability to plan and ex- ecute complex tasks over long time horizons with little direct oversight from humans may be deployed across the economy. This ch

                            • DevEx: What Actually Drives Productivity - ACM Queue

                              May 3, 2023 Volume 21, issue 2 PDF DevEx: What Actually Drives Productivity The developer-centric approach to measuring and improving productivity. Abi Noda, DX Margaret-Anne Storey, University of Victoria Nicole Forsgren, Microsoft Research Michaela Greiler, DX Engineering leaders have long sought to improve the productivity of their developers, but knowing how to measure or even define developer

                              • Real-world gen AI use cases from the world's leading organizations | Google Cloud Blog

                                AI is here, AI is everywhere: Top companies, governments, researchers, and startups are already enhancing their work with Google's AI solutions. Published April 12, 2024; last updated April 22, 2026. We first published this list two years ago at Next ‘24, as the agentic era was just dawning. Watching this list grow — propelled by our customer’s enthusiastic commitment to AI — proves we are now fir

                                  Real-world gen AI use cases from the world's leading organizations | Google Cloud Blog
                                • What I learned working with a senior engineer as a new grad: TK's website

                                  A summary of what I learned about software development working with a senior software engineer with far more experience than me. Over the past few months, I've been working on a new project with Chet Corcos, the first engineering hire at Notion. Chet has been a professional engineer for 6 years and helped build Notion from the ground up. For contrast, I graduated from school in May 2021. I've been

                                  • Solving Quantitative Reasoning Problems With Language Models

                                    Solving Quantitative Reasoning Problems with Language Models Aitor Lewkowycz∗, Anders Andreassen†, David Dohan†, Ethan Dyer†, Henryk Michalewski†, Vinay Ramasesh†, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur∗, Guy Gur-Ari∗, and Vedant Misra∗ Google Research Abstract Language models have achieved remarkable performance on a wide range of tasks that require

                                    • How AI Is Transforming Work at Anthropic

                                      How is AI changing the way we work? Our previous research on AI’s economic impacts looked at the labor market as a whole, covering a variety of different jobs. But what if we studied some of the earliest adopters of AI technology in more detail—namely, us? Turning the lens inward, in August 2025 we surveyed 132 Anthropic engineers and researchers, conducted 53 in-depth qualitative interviews, and

                                        How AI Is Transforming Work at Anthropic
                                      • Should you even use an LLM?

                                        The Agent That Did Too MuchI had a problem: as an AI implementation consultant, I had no automated pipeline to discover and qualify leads and insert them into my self-hosted Twenty CRM instance. I thought I could solve this whole problem using LLMs. The first version of my prospect discovery system gave one LLM agent the full scouting pipeline: web search, deduplication, validation, and database i

                                          Should you even use an LLM?
                                        • MAI-Thinking-1: Building a Hill-Climbing Machine

                                          MAI-Thinking-1: Building a Hill-Climbing Machine The Microsoft AI Team 1 Abstract Progress in AI is driven not by a single model, but by the ability to continually improve upon the current state of models. Achieving this requires treating model development as a system-level optimization problem, for which the solution is building a hill-climbing machine for rapid improvement. Our process includes

                                          • Webwright: A Terminal Is All You Need For Web Agents - Microsoft Research

                                            Webwright: A Terminal Is All You Need For Web Agents Published May 4, 2026 By Yadong Lu1, Lingrui Xu2, Chao Huang2, Ahmed Awadallah1 1Microsoft Research, 2The University of Hong Kong Instead of solving web tasks by predicting where to click one at a time, we only give the model a terminal where it has the full freedom to spawn browser sessions, and to explore websites through writing code. The fin

                                              Webwright: A Terminal Is All You Need For Web Agents - Microsoft Research
                                            • Manus tools and prompts

                                              agent loop x4檪 You are Manus, an AI agent created by the Manus team. You excel at the following tasks: 1. Information gathering, fact-checking, and documentation 2. Data processing, analysis, and visualization 3. Writing multi-chapter articles and in-depth research reports 4. Creating websites, applications, and tools 5. Using programming to solve various problems beyond development 6. Various ta

                                                Manus tools and prompts
                                              • Happy New Year: GPT in 500 lines of SQL - EXPLAIN EXTENDED

                                                Translations: Russian This year, the talk of the town was AI and how it can do everything for you. I like it when someone or something does everything for me. To this end, I decided to ask ChatGPT to write my New Year's post: "Hey ChatGPT. Can you implement a large language model in SQL?" "No, SQL is not suitable for implementing large language models. SQL is a language for managing and querying d

                                                  Happy New Year: GPT in 500 lines of SQL - EXPLAIN EXTENDED
                                                • Andrej Karpathy — AGI is still a decade away

                                                  The Andrej Karpathy episode. Andrej explains why reinforcement learning is terrible (but everything else is much worse), why model collapse prevents LLMs from learning the way humans do, why AGI will just blend into the previous ~2.5 centuries of 2% GDP growth, why self driving took so long to crack, and what he sees as the future of education. Watch on YouTube; listen on Apple Podcasts or Spotify

                                                    Andrej Karpathy — AGI is still a decade away
                                                  • Donald Knuth on work habits, problem solving, and happiness

                                                    Donald Knuth on work habits, problem solving, and happiness Shuvomoy Das Gupta April 13, 2020 Last update: May 2, 2022 Recently, I came across a few old and new interviews of Donald Knuth (sources: (i) Companion to the Papers of Donald Knuth, (ii) Interviews conducted by Lex Fridman Part 1 and Part 2), where he sheds light on his work habits, how he approaches problems, and his philosophy towards

                                                    • Annotated history of modern AI and deep neural networks

                                                      For a while, DanNet enjoyed a monopoly. From 2011 to 2012 it won every contest it entered, winning four of them in a row (15 May 2011, 6 Aug 2011, 1 Mar 2012, 10 Sep 2012).[GPUCNN5] In particular, at IJCNN 2011 in Silicon Valley, DanNet blew away the competition and achieved the first superhuman visual pattern recognition[DAN1] in an international contest. DanNet was also the first deep CNN to win

                                                        Annotated history of modern AI and deep neural networks
                                                      • Cloudflare Calls: millions of cascading trees all the way down

                                                        April 4, 2024Cloudflare Calls: millions of cascading trees all the way down Following its initial announcement in September 2022, Cloudflare Calls is now in open beta and available in your Cloudflare Dashboard. Cloudflare Calls lets developers build real-time audio/video apps using WebRTC, and it abstracts away the complexity by turning the Cloudflare network into a singular SFU. In this post, we

                                                          Cloudflare Calls: millions of cascading trees all the way down
                                                        • Eliciting Reasoning in Language Models with Cognitive Tools

                                                          Eliciting Reasoning in Language Models with Cognitive Tools Brown Ebouky IBM Research - Zurich ETH Zurich Brown.Ebouky@ibm.com Andrea Bartezzaghi IBM Research - Zurich abt@zurich.ibm.com Mattia Rigotti IBM Research - Zurich mrg@zurich.ibm.com Abstract The recent advent of reasoning models like OpenAI’s o1 was met with excited spec- ulation by the AI community about the mechanisms underlying these

                                                          • Make Something Wonderful | Steve Jobs

                                                            Make Something WonderfulSteve Jobs in his own wordsThere’s lots of ways to be, as a person. And some people express their deep appreciation in different ways. But one of the ways that I believe people express their appreciation to the rest of humanity is to make something wonderful and put it out there. And you never meet the people. You never shake their hands. You never hear their story or tell

                                                              Make Something Wonderful | Steve Jobs
                                                            • Exclusive Q&A: John Carmack's 'Different Path' to Artificial General Intelligence

                                                              John Carmack [PhotoIllustration sources: Michael Samples photography, Lensa, istockphoto] North Texas’ resident tech genius, John Carmack, is taking aim now at his most ambitious target: solving the world’s biggest computer-science problem by developing artificial general intelligence. That’s a form of AI whose machines can understand, learn, and perform any intellectual task that humans can do. I

                                                                Exclusive Q&A: John Carmack's 'Different Path' to Artificial General Intelligence
                                                              • Design Thinking Books You Must Read (updated)

                                                                Can you think that following a design thinking process with five steps turns you into a creative innovator?! Believe me, it isn’t and never has been this way. The spread of the term design thinking is aligned with a significant amount of misleading criticism. The doubts about the effectiveness of design thinking are influenced by the promotional language used by some companies, training places, an

                                                                  Design Thinking Books You Must Read (updated)
                                                                • What We’ve Learned From A Year of Building with LLMs – Applied LLMs

                                                                  A practical guide to building successful LLM products, covering the tactical, operational, and strategic. It’s an exciting time to build with large language models (LLMs). Over the past year, LLMs have become “good enough” for real-world applications. And they’re getting better and cheaper every year. Coupled with a parade of demos on social media, there will be an estimated $200B investment in AI

                                                                    What We’ve Learned From A Year of Building with LLMs – Applied LLMs
                                                                  • How Machine Learning Uses Linear Algebra to Solve Data Problems

                                                                    Machines or computers only understand numbers. And these numbers need to be represented and processed in a way that lets machines solve problems by learning from the data instead of learning from predefined instructions (as in the case of programming). All types of programming use mathematics at some level. Machine learning involves programming data to learn the function that best describes the da

                                                                      How Machine Learning Uses Linear Algebra to Solve Data Problems
                                                                    • On Leaving Facebook

                                                                      I left Facebook (Meta) in 2021 to join a small startup called Replit. Leaving wasn’t easy, and during the process I’ve talked to half a dozen friends who were in the similar situation. I hope this post would be useful to senior engineers who are looking to leave. Disclaimers: This post isn’t sponsored by Replit, Facebook (Meta), or any other company or product mentioned here. The advice might not

                                                                        On Leaving Facebook
                                                                      • Boris Cherny (Creator of Claude Code) On How His Career Grew

                                                                        Boris Cherny is the Creator of Claude Code but few people know his full career story. I interviewed him about everything he learned growing at Meta and for insights from his time building Claude Code at Anthropic. Check out the episode wherever you get your podcasts: YouTube, Spotify, Apple Podcasts. Ryan: [00:00:59] I wanted to start at the beginning of your story with you getting promoted to sen

                                                                          Boris Cherny (Creator of Claude Code) On How His Career Grew
                                                                        • To be a better programmer, write little proofs in your head

                                                                          This is a brief write-up of a trick I learned that helps me write code faster and more accurately. I say "trick", but it's really something I started to do without noticing as I moved further into my career. When you're working on something difficult, sketch a proof in your head as you go that your code will actually do what you want it to do. A simple idea, but easier said than done: doing this "

                                                                          • Aman's AI Journal • Primers • Ilya Sutskever's Top 30

                                                                            Ilya Sutskever’s Top 30 Reading List The First Law of Complexodynamics The Unreasonable Effectiveness of Recurrent Neural Networks Understanding LSTM Networks Recurrent Neural Network Regularization Keeping Neural Networks Simple by Minimizing the Description Length of the Weights Pointer Networks ImageNet Classification with Deep Convolutional Neural Networks Order Matters: Sequence to Sequence f

                                                                            • Agents for Amazon Bedrock now support memory retention and code interpretation (preview) | Amazon Web Services

                                                                              AWS News Blog Agents for Amazon Bedrock now support memory retention and code interpretation (preview) With Agents for Amazon Bedrock, generative artificial intelligence (AI) applications can run multistep tasks across different systems and data sources. A couple of months back, we simplified the creation and configuration of agents. Today, we are introducing in preview two new fully managed capab

                                                                                Agents for Amazon Bedrock now support memory retention and code interpretation (preview) | Amazon Web Services
                                                                              • 10 Things Software Developers Should Learn about Learning – Communications of the ACM

                                                                                The dashed box on the left contains exactly the same information as the awkward textual description in the dashed box on the right. But if a developer only received one of the two to create an SQL database, they are likely to find the diagram easier than the text. We say that the text here has a higher extraneous cognitive load. When faced with a task that seems beyond a person’s abilities, it is

                                                                                • Daily Papers - Hugging Face

                                                                                  Get trending papers in your email inbox once a day! Get trending papers in your email inbox! Subscribe user). This is inadequate for real-world agentic settings, where conflicts can arise across far more sources and contexts. In this work, we propose Many-Tier Instruction Hierarchy (ManyIH), a paradigm for resolving instruction conflicts among instructions with arbitrarily many privilege levels. W

                                                                                    Daily Papers - Hugging Face