24小时 Together AI 实时更新
Together AI 模型服务动态
LATEST UPDATES
最新获取
2026-08-16
-
Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets
How Together served MiniMax-M3 efficiently with KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.
Together AI产品动态查看详情 -
Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification
Together AI has earned ISO 27001:2022 certification, validating our commitment to enterprise-grade security for production AI workloads.
Together AI产品动态查看详情 -
Kimi K2.7 Code vs Claude Fable 5: Landing pages that cost 94% less
We generated 12 landing pages with Kimi K2.7 Code and Claude Fable 5. Kimi cost 94% less and scored within a few points on every page. Here's what actually moved the needle.
Together AI产品动态查看详情 -
ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)
ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.
Together AI产品动态查看详情 -
Together AI at ICML 2026: frontier research across the full stack
Nine papers at ICML 2026 across the full stack. The research that becomes the Together platform. Find us at booth B714 in Seoul.
Together AI产品动态查看详情 -
Announcing our $800M Series C to accelerate the shift to open-source AI
We raised $800M to accelerate the shift to open-source AI. Here's why the economics of closed models don't scale, and what we're building next.
Together AI产品动态查看详情 -
Open, convenient and predictable: Introducing Provisioned Throughput
Provisioned Throughput gives you reserved inference capacity for frontier open models like MiniMax M3 and GLM-5.2. Token-based pricing, a 99% uptime SLA, and up to 90% lower cost than proprietary APIs. No GPU-hour math, no infrastructure to manage.
Together AI产品动态查看详情 -
New in Together GPU Clusters: Reliability and control for production GPU clusters
See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.
Together AI产品动态查看详情 -
Together AI brings Thinking Machines Lab’s new model Inkling on day 0
Together AI offers day zero access to Inkling, Thinking Machines Lab's multimodal mixture-of-experts model for text, image, and audio reasoning.
Together AI产品动态查看详情 -
What does 99.9% uptime mean for inference?
Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference provider before you commit.
Together AI产品动态查看详情 -
Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community
No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.
Together AI产品动态查看详情 -
The production platform for open-weight AI inference
Run open models in production with full control over performance, cost, and quality. Deploy in minutes, roll out safely, and scale to your SLOs.
Together AI产品动态查看详情 -
Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding
We ran 452 DeepSWE rollouts on Kimi K3 and Claude Fable 5. Fable leads pass@1 by 1.4 points; Kimi K3 wins pass@4 and delivers 2.8x the solves per dollar.
Together AI产品动态查看详情 -
Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
We ran 904 DeepSWE rollouts on Kimi K3 and GPT-5.6 Sol. Sol leads pass@1; Kimi K3 wins pass@4 at 2.8x the solves per dollar, and routing between them reaches ~85.6%.
Together AI产品动态查看详情 -
ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale
ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.
Together AI产品动态查看详情 -
Configuring Dedicated Model Inference
The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together.
Together AI产品动态查看详情 -
Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
Together AI partners with Moonshot AI to natively serve Kimi models, starting with the 2.8T parameter Kimi K3, with day zero access and post-training.
Together AI产品动态查看详情 -
Autoscaling endpoints for LLM inference
GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.
Together AI产品动态查看详情 -
Kimi K3: The Complete Developer Guide
Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.
Together AI产品动态查看详情 -
DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding
We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
Together AI产品动态查看详情