首页 / AI每日快讯

24小时 Together AI 实时更新

Together AI 模型服务动态

LATEST UPDATES

最新获取

Together AI · 46 条快讯
全部更新按发布时间倒序

2026-09-02

  1. GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

    We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.

    Together AI产品动态查看详情
  2. GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

    We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.

    Together AI产品动态查看详情
  3. GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

    We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

    Together AI产品动态查看详情
  4. DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

    We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

    Together AI产品动态查看详情
  5. DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

    We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.

    Together AI产品动态查看详情
  6. A/B test models in production

    Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.

    Together AI产品动态查看详情
  7. Capacity without conflict: A guide to multi-tenant GPU cluster design for AI-native teams

    Learn how AI-native companies design multi-tenant GPU clusters that pool capacity without sacrificing team isolation — and how Together AI makes it work in practice.

    Together AI产品动态查看详情
  8. Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding

    Rollout is the silent bottleneck in RL post-training. DAS fixes it with adaptive speculative decoding — up to 50% faster, zero degradation in reward quality.

    Together AI产品动态查看详情
  9. Together AI Brings NVIDIA Nemotron 3 Nano Omni to Developers on Day 0

    NVIDIA Nemotron 3 Nano Omni is now on Together AI: a single open model that reasons across video, images, audio, and text, built for agentic workloads at scale.

    Together AI产品动态查看详情
  10. DeepSeek-V4 Pro now available on Together AI

    DeepSeek-V4 Pro is now available on Together AI with 512K context, controllable reasoning modes, and cached-input pricing for long-context reasoning workloads like code agents, document intelligence, and research synthesis.

    Together AI产品动态查看详情
  11. Announcing Together AI and Adaption Partnership

    Together AI and Adaption partner to bring Together Fine-Tuning natively into Adaptive Data, helping teams optimize datasets, run fine-tuning, evaluate results, and deploy stronger open models.

    Together AI产品动态查看详情
  12. From 732 bytes to nowhere: shutting down Copy Fail in production

    How Together AI responded to the Copy Fail Linux kernel bug (CVE-2026-31431): disabling the affected crypto interface fleet wide and patching safely.

    Together AI产品动态查看详情
  13. Foundational research powering efficient inference at scale

    As AI moves from research to production, the challenge for AI-native teams shifts from building models to running them — efficiently, reliably, and at scale.

    Together AI产品动态查看详情
  14. Deploy and inference any model from HuggingFace

    Learn how to deploy any Hugging Face model in one session using Goose and Together's Dedicated Container Inference. Skip the setup complexity — one prompt gets your model running in a production-grade GPU environment on release day.

    Together AI产品动态查看详情
  15. Serving DeepSeek-V4: why million-token context is an inference systems problem

    DeepSeek-V4 makes million-token context a serving-systems problem. Together AI explores the inference work behind V4 on NVIDIA HGX B200, including compressed KV layouts, prefix caching, kernel maturity, and endpoint profiles for long-context workloads.

    Together AI产品动态查看详情
  16. Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices

    Voice finder helps developers search, match, filter, and audition 600+ voices across Together AI TTS models using natural-language prompts or uploaded audio samples.

    Together AI产品动态查看详情
  17. Violin: An open-source video translation skill that breaks language barriers

    Violin is an open-source AI video translation tool that combines speech recognition, LLM translation, and text-to-speech to make video content accessible across languages.

    Together AI产品动态查看详情
  18. Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference

    Together AI partners with Pearl Research Labs to launch a discounted Pearl-powered inference endpoint for Gemma-4-31B-it-pearl, using Proof of Useful Work to turn AI workloads into crypto emissions.

    Together AI产品动态查看详情
  19. Benchmarking inference at scale: coding agents

    Real-world inference benchmarks for coding agents: 31% more TPS than TensorRT-LLM, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4.6.

    Together AI产品动态查看详情
  20. How Together AI built the world’s fastest speech-to-text stack

    Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.

    Together AI产品动态查看详情