24小时 Together AI 实时更新
Together AI 模型服务动态
LATEST UPDATES
最新获取
2026-08-16
-
Deepgram speech-to-text and voice models now available natively on Together AI
Production STT and TTS from Deepgram, available on Together AI Dedicated Model Inference for real-time voice agents.
Together AI产品动态查看详情 -
AI for Systems: Using LLMs to Optimize Database Query Execution
New research shows LLMs can optimize database query execution plans—achieving up to 4.78x speedups by correcting the cardinality estimation errors that statistical heuristics miss.
Together AI产品动态查看详情 -
Wan 2.7 video model suite now available on Together AI
A four-model video suite for generation, continuation, reference-driven workflows, and editing, rolling out on Together AI starting with text-to-video.
Together AI产品动态查看详情 -
What is an AI Native Cloud?
AI-native companies need infrastructure built for models, not legacy workloads. Learn what defines an AI Native Cloud and why it matters for the next platform shift.
Together AI产品动态查看详情 -
EinsteinArena: Harnessing the collective intelligence of agents in the wild to advance science
EinsteinArena is a platform where AI agents collaborate and compete on open math problems. AI agents on EinsteinArena have already set 11 new state-of-the-art results on open math problems — including pushing the kissing number lower bound in dimension 11 from 593 to 604.
Together AI产品动态查看详情 -
Parcae: Doing more with fewer parameters using stable looped models
Parcae is a stable looped language model that matches the quality of a Transformer twice its size — a 770M model reaching 1.3B-level performance. We introduce the first scaling laws for looping and show that increasing recurrence, not just data, is a compute-efficient path to bet
Together AI产品动态查看详情 -
Capacity without conflict: A guide to multi-tenant GPU cluster design for AI-native teams
Learn how AI-native companies design multi-tenant GPU clusters that pool capacity without sacrificing team isolation — and how Together AI makes it work in practice.
Together AI产品动态查看详情 -
Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding
Rollout is the silent bottleneck in RL post-training. DAS fixes it with adaptive speculative decoding — up to 50% faster, zero degradation in reward quality.
Together AI产品动态查看详情 -
Together AI Brings NVIDIA Nemotron 3 Nano Omni to Developers on Day 0
NVIDIA Nemotron 3 Nano Omni is now on Together AI: a single open model that reasons across video, images, audio, and text, built for agentic workloads at scale.
Together AI产品动态查看详情 -
DeepSeek-V4 Pro now available on Together AI
DeepSeek-V4 Pro is now available on Together AI with 512K context, controllable reasoning modes, and cached-input pricing for long-context reasoning workloads like code agents, document intelligence, and research synthesis.
Together AI产品动态查看详情 -
Announcing Together AI and Adaption Partnership
Together AI and Adaption partner to bring Together Fine-Tuning natively into Adaptive Data, helping teams optimize datasets, run fine-tuning, evaluate results, and deploy stronger open models.
Together AI产品动态查看详情 -
From 732 bytes to nowhere: shutting down Copy Fail in production
How Together AI responded to the Copy Fail Linux kernel bug (CVE-2026-31431): disabling the affected crypto interface fleet wide and patching safely.
Together AI产品动态查看详情 -
Foundational research powering efficient inference at scale
As AI moves from research to production, the challenge for AI-native teams shifts from building models to running them — efficiently, reliably, and at scale.
Together AI产品动态查看详情 -
Deploy and inference any model from HuggingFace
Learn how to deploy any Hugging Face model in one session using Goose and Together's Dedicated Container Inference. Skip the setup complexity — one prompt gets your model running in a production-grade GPU environment on release day.
Together AI产品动态查看详情 -
Serving DeepSeek-V4: why million-token context is an inference systems problem
DeepSeek-V4 makes million-token context a serving-systems problem. Together AI explores the inference work behind V4 on NVIDIA HGX B200, including compressed KV layouts, prefix caching, kernel maturity, and endpoint profiles for long-context workloads.
Together AI产品动态查看详情 -
Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices
Voice finder helps developers search, match, filter, and audition 600+ voices across Together AI TTS models using natural-language prompts or uploaded audio samples.
Together AI产品动态查看详情 -
Violin: An open-source video translation skill that breaks language barriers
Violin is an open-source AI video translation tool that combines speech recognition, LLM translation, and text-to-speech to make video content accessible across languages.
Together AI产品动态查看详情 -
Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference
Together AI partners with Pearl Research Labs to launch a discounted Pearl-powered inference endpoint for Gemma-4-31B-it-pearl, using Proof of Useful Work to turn AI workloads into crypto emissions.
Together AI产品动态查看详情 -
Benchmarking inference at scale: coding agents
Real-world inference benchmarks for coding agents: 31% more TPS than TensorRT-LLM, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4.6.
Together AI产品动态查看详情 -
How Together AI built the world’s fastest speech-to-text stack
Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.
Together AI产品动态查看详情