並び順

ブックマーク数

期間指定

  • から
  • まで

1 - 26 件 / 26件

新着順 人気順

https captions ai apiの検索結果1 - 26 件 / 26件

  • SpotifyのPodcast、OpenAIの技術で本人の声での多言語吹き替えが可能に

    Spotifyは「クリエイター自身の声を使うことで、音声翻訳はこれまで以上にリアルな方法で世界中のリスナーにホストのインスピレーションを受け取る力を与える」と語った。 ダニエル・エクCEOのXのポストで、スティーブン・バートレット氏とレックス・フリードマン氏のスペイン語吹き替えを試聴できる。 関連記事 ChatGPT、“目”と“耳”の実装を発表 写真の内容を認識、発話機能でおしゃべりも可能に 米OpenAIのチャットAI「ChatGPT」に、画像認識、音声認識、発話機能が搭載された。今後2週間かけて、PlusユーザーとEnterpriseユーザーに展開するという。 YouTube、クリエイター向けイベントでAI搭載の複数ツールを発表 YouTubeはクリエイター向けイベントを開催し、複数の編集ツールを発表した。YouTubeショートの背景を生成AIで作る「Dream Screen」など、A

      SpotifyのPodcast、OpenAIの技術で本人の声での多言語吹き替えが可能に
    • 「公開するApple vs. 隠すOpenAI」アップルが300億パラメータのマルチモーダルAI「MM1」発表。重要論文5本を解説(生成AIウィークリー) | テクノエッジ TechnoEdge

      1週間分の生成AI関連論文の中から重要なものをピックアップし、解説をする連載です。第38回目は、生成AI最新論文の概要5つを紹介します。 Appleが最大300億パラメータを持つマルチモーダル大規模言語モデル「MM1」を開発> OpenAIなどのクローズド大規模言語モデルの一部を許可なく取得する攻撃、Googleなどが開発 GPT-3.5の隠れ層のサイズを約4096と推定。非公開LLMの中身を抽出する手法 WebページのスクリーンショットからHTMLコードを生成するAIモデル「Sightseer」をHugging Faceが開発 実世界に強いマルチモーダル大規模言語モデル「DeepSeek-VL」 Appleが最大300億パラメータを持つマルチモーダル大規模言語モデル「MM1」を開発既存のマルチモーダル大規模言語モデル(MLLM)は、透明性の観点からクローズドモデルとオープンモデルの2つに

        「公開するApple vs. 隠すOpenAI」アップルが300億パラメータのマルチモーダルAI「MM1」発表。重要論文5本を解説(生成AIウィークリー) | テクノエッジ TechnoEdge
      • 「Illustrious」はなぜ強いのか?次世代AIイラストモデルの論文を日本語で読む|kazumu(シズミ)

        原著論文タイトル:Illustrious: an Open Advanced Illustration Model(2024年9月30日に発表) 著者:Onoma Research arXivリンク:https://arxiv.org/abs/2409.19946 ライセンス:本論文は Creative Commons Attribution 4.0 International (CC BY 4.0) ライセンスのもとで公開されています。 本翻訳も同ライセンスに準拠しており、著作者のクレジットを保持したうえで自由に再利用・再配布・改変が可能です。 Illustrious:オープンで高度なイラスト生成モデルサン・ヒョン・パク(Sang Hyun Park)、コ・ジュンヨン(Jun Young Koh)、イ・ジュンハ(Junha Lee)、ジョイ・ソン(Joy Song)、 キム・ドンハ(Do

          「Illustrious」はなぜ強いのか?次世代AIイラストモデルの論文を日本語で読む|kazumu(シズミ)
        • 誰でもわかるStable Diffusion Kohya_ssを使ったLoRA学習設定を徹底解説 - 人工知能と親しくなるブログ

          前回の記事では、Stable Diffusionモデルを追加学習するためのWebUI環境「kohya_ss」の導入法について解説しました。 今回は、LoRAのしくみを大まかに説明し、その後にkohya_ssを使ったLoRA学習設定について解説していきます。 ※今回の記事は非常に長いです! この記事では「各設定の意味」のみ解説しています。 「学習画像の用意のしかた」とか「画像にどうキャプションをつけるか」とか「どう学習を実行するか」は解説していません。学習の実行法についてはまた別の記事で解説したいと思います。 LoRAの仕組みを知ろう 「モデル」とは LoRAは小さいニューラルネットを追加する 小さいニューラルネットの構造 LoRA学習対象1:U-Net RoLA学習対象2:テキストエンコーダー kohya_ssを立ち上げてみよう LoRA学習の各設定 LoRA設定のセーブ、ロード Sour

            誰でもわかるStable Diffusion Kohya_ssを使ったLoRA学習設定を徹底解説 - 人工知能と親しくなるブログ
          • 日本語CLIP 学習済みモデルと評価用データセットの公開

            はじめに 基盤モデル がAIの新潮流となりました。基盤モデルというとやはり大規模言語モデルが人気ですが、リクルートでは、画像を扱えるモデルの開発にも注力しています。画像を扱える基盤モデルの中でも代表的なモデルのCLIPは実務や研究のさまざまな場面で利用されています。CLIPの中には日本語に対応したものも既に公開されていますが、その性能には向上の余地がある可能性があると私たちは考え、仮説検証を行ってきました。今回はその検証の過程で作成したモデルと評価用データセットの公開をしたいと思います。 公開はHugging Face上で行っていますが、それに合わせて本記事では公開されるモデルやデータセットの詳細や、公開用モデルの学習の工夫などについて紹介します。 本記事の前半では、今回公開するモデルの性能や評価用データセットの内訳、学習の設定について紹介します。記事の後半では大規模な学習を効率的に実施す

              日本語CLIP 学習済みモデルと評価用データセットの公開
            • いよいよ「ノートPCだけでチャットAIが動く」時代がやってきた - ライブドアニュース

              Photo: 小野寺しんいち OpenAIが、ローカルで動くAI、gpt-ossを8月に発表しました。これは、業界がローカルAI(オンデバイスAI)へ舵を切った決定的なニュースだったと思います 日本HPの岡戸伸樹社長は、オンデバイスAI時代の幕開けをこう宣言します。 Photo: 小野寺しんいち 岡戸伸樹氏 いまや仕事でも日常でも、やGeminiのようなAIを使うのが当たり前になってきたという人も多いんじゃないでしょうか。 でも、AI業界の進化は恐ろしいほど早い。私たちが日常的に使えるAIは、次なるステージへと移行しようとしています。それが、「オンデバイスAI」です。 2025年10月3日に、日本HPが開催した「」では、オンデバイスAIの最前線が語られました。 本記事では、カンファレンスで得た情報をもとに、「結局オンデバイスAIって何がすごいの?」、「AI PCって何ができるの?」を解説し

                いよいよ「ノートPCだけでチャットAIが動く」時代がやってきた - ライブドアニュース
              • Opus 4.5 is going to change everything

                Edit: A lot of folks have been asking what worfklows I used to write these apps. I used GitHub Copilot in VS Code with a custom agent prompt that you’ll find toward the end of this post. Context7 was the only MCP I used. I mostly just used the built-in voice dictation feature and talked to Claude. No fancy workflows, planning, etc required. The agent harness in VS Code for Opus 4.5 is so good - yo

                  Opus 4.5 is going to change everything
                • Unlock a new era of innovation with Windows Copilot Runtime and Copilot+ PCs

                  Unlock a new era of innovation with Windows Copilot Runtime and Copilot+ PCs I am excited to be back at Build with the developer community this year. Over the last year, we have worked on reimagining  Windows PCs and yesterday, we introduced the world to a new category of Windows PCs called Copilot+ PCs. Copilot+ PCs are the fastest, most intelligent Windows PCs ever with AI infused at every layer

                    Unlock a new era of innovation with Windows Copilot Runtime and Copilot+ PCs
                  • Real-world gen AI use cases from the world's leading organizations | Google Cloud Blog

                    AI is here, AI is everywhere: Top companies, governments, researchers, and startups are already enhancing their work with Google's AI solutions. Published April 12, 2024; last updated April 22, 2026. We first published this list two years ago at Next ‘24, as the agentic era was just dawning. Watching this list grow — propelled by our customer’s enthusiastic commitment to AI — proves we are now fir

                      Real-world gen AI use cases from the world's leading organizations | Google Cloud Blog
                    • 🎙️ MacWhisper

                      Quickly and easily transcribe audio files into text with OpenAI's state-of-the-art transcription technology Whisper as well as Nvidia Parakeet. Whether you're recording a meeting, lecture, or other important audio, MacWhisper quickly and accurately transcribes your audio files into text. 📲 MacWhisper is now also available on iPhone and iPad, download it here. 🎁 Get 5 euros off in January by clic

                        🎙️ MacWhisper
                      • SceneXplain - Leading AI Solution for Image Captions and Video Summaries

                        Experience cutting-edge computer vision with our premier image captioning and video summarization algorithms. Tailored for content creators, media professionals, SEO experts, and e-commerce enterprises. Featuring multilingual support and seamless API integration. Elevate your digital presence today.

                          SceneXplain - Leading AI Solution for Image Captions and Video Summaries
                        • いよいよ「ノートPCだけでチャットAIが動く」時代がやってきた | ギズモード・ジャパン

                          OpenAIが、ローカルで動くAI、gpt-ossを8月に発表しました。これは、業界がローカルAI(オンデバイスAI)へ舵を切った決定的なニュースだったと思います 日本HPの岡戸伸樹社長は、オンデバイスAI時代の幕開けをこう宣言します。 Photo: 小野寺しんいち岡戸伸樹氏いまや仕事でも日常でも、ChatGPTやGeminiのようなAIを使うのが当たり前になってきたという人も多いんじゃないでしょうか。 でも、AI業界の進化は恐ろしいほど早い。私たちが日常的に使えるAIは、次なるステージへと移行しようとしています。それが、「オンデバイスAI」です。 2025年10月3日に、日本HPが開催した「HP Future of Work AI Conference 2025」では、オンデバイスAIの最前線が語られました。 本記事では、カンファレンスで得た情報をもとに、「結局オンデバイスAIって何がす

                            いよいよ「ノートPCだけでチャットAIが動く」時代がやってきた | ギズモード・ジャパン
                          • Build 2026: Furthering Windows as the trusted platform for development

                            Build 2026: Furthering Windows as the trusted platform for development Build is one of our favorite moments each year – a chance to connect with the global developer community and share what we’ve been building. Over the past year, we have connected with many developers pushing the boundaries of what’s possible on Windows. What we consistently hear is that you want a platform that meets you where

                              Build 2026: Furthering Windows as the trusted platform for development
                            • stability-ai/stable-diffusion | Run with an API on Replicate

                              Run time and cost This model costs approximately $0.0037 to run on Replicate, or 270 runs per $1, but this varies depending on your inputs. It is also open source and you can run it on your own computer with Docker. This model runs on Nvidia A100 (80GB) GPU hardware. Predictions typically complete within 3 seconds. Stable Diffusion is a latent text-to-image diffusion model capable of generating ph

                                stability-ai/stable-diffusion | Run with an API on Replicate
                              • Fuyu-8B: A Multimodal Architecture for AI Agents

                                Today, we’re releasing Fuyu-8B with an open license (CC-BY-NC)—we’re excited to see what the community builds on top of it! We also discuss results for Fuyu-Medium (a larger model we’re not releasing) and provide a sneak peek of some capabilities that are exclusive to our internal models. Because this is a raw model release, we have not added further instruction-tuning, postprocessing or sampling

                                  Fuyu-8B: A Multimodal Architecture for AI Agents
                                • Building A Generative AI Platform

                                  After studying how companies deploy generative AI applications, I noticed many similarities in their platforms. This post outlines the common components of a generative AI platform, what they do, and how they are implemented. I try my best to keep the architecture general, but certain applications might deviate. This is what the overall architecture looks like. This is a pretty complex system. Thi

                                    Building A Generative AI Platform
                                  • DOMDOMタイムス#15: canvas-based renderingとa11y。いま、そしてこれから

                                    DOMDOMタイムス#15: canvas-based renderingとa11y。いま、そしてこれから はい、DOMDOMタイムスです。知ってるよって?この挨拶は、まあ準備体操みたいなものなので。いつだって欠かさずやっていきます👶 さて、今日はcanvas-based renderingとa11yの話です。この前のJSConfセッションの最後のあたりで話したことと重複が大きいですが、面白がってくれる方が多かったので文章にしておこうと思いました。 (一応JSConfセッションの資料へのリンクも載せておきますね) canvas-based renderingの波 canvas-based renderingという言葉は、この記事では「canvasをゴリゴリに使ってwebコンテンツをレンダリングすること」ってくらいの意味で使います。通常のform要素やdiv要素ではなくて、canvas要素

                                      DOMDOMタイムス#15: canvas-based renderingとa11y。いま、そしてこれから
                                    • Power video semantic search with Amazon Nova Multimodal Embeddings | Amazon Web Services

                                      Artificial Intelligence Power video semantic search with Amazon Nova Multimodal Embeddings Video semantic search is unlocking new value across industries. The demand for video-first experiences is reshaping how organizations deliver content, and customers expect fast, accurate access to specific moments within video. For example, sports broadcasters need to surface the exact moment a player scored

                                        Power video semantic search with Amazon Nova Multimodal Embeddings | Amazon Web Services
                                      • GitHub - ronak-create/FableCut: Zero-dependency browser video editor that AI agents can drive — JSON timeline, MCP + REST, live-reloading UI

                                        Most "AI video" tools hide the edit behind an API. FableCut flips that: the project file is the interface. project.json describes media, clips, tracks, effects, keyframes and transitions — any process that can write JSON can edit video, and the open browser UI hot-reloads within ~150 ms via server-sent events. A human and an agent can work on the same timeline at the same time. Editing 4 video tra

                                          GitHub - ronak-create/FableCut: Zero-dependency browser video editor that AI agents can drive — JSON timeline, MCP + REST, live-reloading UI
                                        • 【2021年版】SNS主要アップデート情報を総ざらい! 年末年始の総復習に « 株式会社ガイアックス

                                          a]:flex [&>a]:flex-row [&>a]:justify-between [&>a]:py-[18px] [&>a]:border-t [&>a]:border-lightgray [&>a]:border-opacity-20 [&_li]:my-1 [&_li]:list-['-_'] [&_li]:py-[18px] [&_li]:border-t [&_li]:border-lightgray [&_li]:border-opacity-20 [&_.Label]:transition-all [&_.Label]:w-fit [&_.content]:transition-all [&_.content]:h-0 [&_.content]:pt-0 [&_.content]:px-5 [&_.content]:overflow-hidden [&_.toggle:

                                            【2021年版】SNS主要アップデート情報を総ざらい! 年末年始の総復習に « 株式会社ガイアックス
                                          • A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT

                                            111 A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT YIHAN CAO∗, Lehigh University & Carnegie Mellon University, USA SIYU LI, Lehigh University, USA YIXIN LIU, Lehigh University, USA ZHILING YAN, Lehigh University, USA YUTONG DAI, Lehigh University, USA PHILIP S. YU, University of Illinois at Chicago, USA LICHAO SUN, Lehigh University, USA Recen

                                            • 講義動画における生成 AI を活用した字幕生成 - スタディサプリ Product Team Blog

                                              こんにちは、『スタディサプリ』の iOS エンジニアのヴァンサンです。 先日、『スタディサプリ』の一部の講座の動画に日本語字幕が追加されました。音声と同じ言語の字幕は、聴覚に障がいのあるユーザーだけでなく、音声が聞こえづらい環境や、イヤホンが手元になく音を出せない環境でも有用です。さらに、字幕データ自体も検索や内容のまとめなど、さまざまな用途での活用が期待できます。そのデータがなければ、せっかく制作したコンテンツをフル活用できないでしょう。 この記事では、私たちが自動生成を選んだ経緯や字幕生成のプロセスを紹介します。私の生成 AI に関する知識はまだ浅く、改善の余地は多分にあります。また、AI 技術は急速に進化しているため、ここで紹介する方法はすぐに時代遅れになる可能性もあります。それでも、この取り組みが何かの参考になれば幸いです。 字幕 まず、生成について説明する前に、字幕の基本的な概念

                                                講義動画における生成 AI を活用した字幕生成 - スタディサプリ Product Team Blog
                                              • Indexing a year of video locally on a 5-year-old M1 Max with Gemma 4 31B

                                                While I slept, my 5-year-old MacBook ran Gemma 4 locally and indexed a year of video May 21, 2026 I'm in the Maasai Mara about half the year, in three-month stretches. Animals out the front of the lodge, motorcycles, friends in the Maasai villages, kids who think a drone is the funniest thing they have ever seen. That's one half of my year. The other half is sixteen-hour days in front of a termina

                                                • GitHub - ComfyUI-Workflow/awesome-comfyui: A collection of awesome custom nodes for ComfyUI

                                                  ComfyUI-Gemini_Flash_2.0_Exp (⭐+172): A ComfyUI custom node that integrates Google's Gemini Flash 2.0 Experimental model, enabling multimodal analysis of text, images, video frames, and audio directly within ComfyUI workflows. ComfyUI-ACE_Plus (⭐+115): Custom nodes for various visual generation and editing tasks using ACE_Plus FFT Model. ComfyUI-Manager (⭐+113): ComfyUI-Manager itself is also a cu

                                                    GitHub - ComfyUI-Workflow/awesome-comfyui: A collection of awesome custom nodes for ComfyUI
                                                  • LAION-400-MILLION OPEN DATASET | LAION

                                                    LAION-400-MILLION OPEN DATASETby: Christoph Schuhmann, 20 Aug, 2021 We present LAION-400M: 400M English (image, text) pairs - see also our Data Centric AI NeurIPS Workshop 2021 paper Concept and Content The LAION-400M dataset is entirely openly, freely accessible. WARNING: be aware that this large-scale dataset is non-curated. It was built for research purposes to enable testing model training on

                                                      LAION-400-MILLION OPEN DATASET | LAION
                                                    • Instagram APIでハッシュタグを収集しダッシュボードを作る方法

                                                      ハッシュタグデータの取得方法 ここからはPythonでの実際のデータ取得方法を解説していきます。 ※スプレッドシートに出力するために最初からGoogle Apps Scriptを利用してしまうのが一番効率的かと思います。 1.検索したいハッシュタグのIDを探す import requests # あらかじめID等は取得しておく instragramID = "xxxxxx" ACCESS_TOKEN = "xxxxxx" # 検索したいワード query = "マキアート" id_search_url = "https://graph.facebook.com/ig_hashtag_search?user_id=" + instragramID + "&q=" + query + "&access_token=" + ACCESS_TOKEN response = requests.get

                                                        Instagram APIでハッシュタグを収集しダッシュボードを作る方法
                                                      1