首页 / AI每日快讯

24小时 AI快讯 实时更新

聚合全球 AI 厂商、模型、API、开源项目与服务状态动态。

LATEST UPDATES

最新获取

全量厂商 · 13722 条快讯
全部更新按发布时间倒序

2026-09-02

  1. From 732 bytes to nowhere: shutting down Copy Fail in production

    How Together AI responded to the Copy Fail Linux kernel bug (CVE-2026-31431): disabling the affected crypto interface fleet wide and patching safely.

    Together AI产品动态查看详情
  2. Foundational research powering efficient inference at scale

    As AI moves from research to production, the challenge for AI-native teams shifts from building models to running them — efficiently, reliably, and at scale.

    Together AI产品动态查看详情
  3. Deploy and inference any model from HuggingFace

    Learn how to deploy any Hugging Face model in one session using Goose and Together's Dedicated Container Inference. Skip the setup complexity — one prompt gets your model running in a production-grade GPU environment on release day.

    Together AI产品动态查看详情
  4. Serving DeepSeek-V4: why million-token context is an inference systems problem

    DeepSeek-V4 makes million-token context a serving-systems problem. Together AI explores the inference work behind V4 on NVIDIA HGX B200, including compressed KV layouts, prefix caching, kernel maturity, and endpoint profiles for long-context workloads.

    Together AI产品动态查看详情
  5. Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices

    Voice finder helps developers search, match, filter, and audition 600+ voices across Together AI TTS models using natural-language prompts or uploaded audio samples.

    Together AI产品动态查看详情
  6. Violin: An open-source video translation skill that breaks language barriers

    Violin is an open-source AI video translation tool that combines speech recognition, LLM translation, and text-to-speech to make video content accessible across languages.

    Together AI产品动态查看详情
  7. Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference

    Together AI partners with Pearl Research Labs to launch a discounted Pearl-powered inference endpoint for Gemma-4-31B-it-pearl, using Proof of Useful Work to turn AI workloads into crypto emissions.

    Together AI产品动态查看详情
  8. Benchmarking inference at scale: coding agents

    Real-world inference benchmarks for coding agents: 31% more TPS than TensorRT-LLM, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4.6.

    Together AI产品动态查看详情
  9. How Together AI built the world’s fastest speech-to-text stack

    Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.

    Together AI产品动态查看详情
  10. Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets

    How Together served MiniMax-M3 efficiently with KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.

    Together AI产品动态查看详情
  11. Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification

    Together AI has earned ISO 27001:2022 certification, validating our commitment to enterprise-grade security for production AI workloads.

    Together AI产品动态查看详情
  12. Kimi K2.7 Code vs Claude Fable 5: Landing pages that cost 94% less

    We generated 12 landing pages with Kimi K2.7 Code and Claude Fable 5. Kimi cost 94% less and scored within a few points on every page. Here's what actually moved the needle.

    Together AI产品动态查看详情
  13. ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)

    ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.

    Together AI产品动态查看详情
  14. Together AI at ICML 2026: frontier research across the full stack

    Nine papers at ICML 2026 across the full stack. The research that becomes the Together platform. Find us at booth B714 in Seoul.

    Together AI产品动态查看详情
  15. Announcing our $800M Series C to accelerate the shift to open-source AI

    We raised $800M to accelerate the shift to open-source AI. Here's why the economics of closed models don't scale, and what we're building next.

    Together AI产品动态查看详情
  16. Open, convenient and predictable: Introducing Provisioned Throughput

    Provisioned Throughput gives you reserved inference capacity for frontier open models like MiniMax M3 and GLM-5.2. Token-based pricing, a 99% uptime SLA, and up to 90% lower cost than proprietary APIs. No GPU-hour math, no infrastructure to manage.

    Together AI产品动态查看详情
  17. New in Together GPU Clusters: Reliability and control for production GPU clusters

    See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.

    Together AI产品动态查看详情
  18. Together AI brings Thinking Machines Lab’s new model Inkling on day 0

    Together AI offers day zero access to Inkling, Thinking Machines Lab's multimodal mixture-of-experts model for text, image, and audio reasoning.

    Together AI产品动态查看详情
  19. What does 99.9% uptime mean for inference?

    Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference provider before you commit.

    Together AI产品动态查看详情
  20. Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community

    No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.

    Together AI产品动态查看详情