TL;DR
이번 기간에는 GPT-6 Astra의 API·ChatGPT 구독자 대상 확대와 Terminal-Bench-Science 점수 64.6%가 함께 언급되며 컴퓨터 사용과 과학 추론 성능이 새 기준으로 부상했습니다. Cybercab에는 Starlink V5가 직접 통합돼 통신 품질이 낮은 장소에서도 차량 운용과 승객용 4K 라이브 영상을 지원하는 구상이 제시됐습니다. 오픈 모델 진영에서는 K2 Horizon의 코드·학습 데이터·레시피·체크포인트 공개와 Ling-3.0-flash-Fin의 금융 AI 활용이 주목받았고, TorchSpec와 vLLM은 speculative decoding용 초안 모델의 학습·서빙 연결을 강화했습니다. 실무 도구에서는 Qwen3.8-Flash 기반 QwenWork의 자동 자료 정리와 deck 생성, Claude Managed Agent 환경을 파일로 동기화하는 ant apply가 등장했습니다. 추론 인프라에서는 PyTorch AOTI·KV cache 조합의 최대 2.38배 속도 향상과 대규모 학습 작업의 fault tolerance 개선이 이어졌습니다.
𝕏 실시간 트렌드 토픽
🔥 GPT-6 Astra의 배포 확대와 Terminal-Bench-Science 성능포스트 5
GPT-6 Astra가 ChatGPT Pro부터 API 고객과 구독자에게 순차 확대될 예정이며, Terminal-Bench-Science에서 GPT-5.6 Sol의 22.4%보다 높은 64.6%가 인용됐습니다.
세부 내용 보기
- GPT-6 Astra 출시와 함께 컴퓨터에서 수행할 수 있는 작업을 처리한다는 설명, ChatGPT Pro 우선 접근, API 고객과 ChatGPT 구독자 대상의 광범위한 배포 계획이 한 흐름으로 묶였습니다. OpenAI 측은 초기 출시가 매끄럽지 못했다며 오류를 바로잡은 뒤 확대하겠다고 밝혔습니다.
- Terminal-Bench-Science 0.1은 과학 영역 작업을 평가하는 기준으로 제시됐고, GPT-6 Astra의 점수는 64.6%로 GPT-5.6 Sol의 22.4%보다 42.2%포인트 높았습니다. 같은 벤치마크가 GPT-6 Astra 출시 당일 주요 평가 항목으로 올라오면서 모델 공개와 외부 평가가 연결됐습니다.
- GPT-6 Astra의 System Card에는 critical-level cyber capability, 개선된 refusals, 낮아진 CoT monitorability가 함께 적혔습니다. 성능 확대와 안전성·관찰 가능성의 변화가 동일한 배포 문서에서 다뤄진 점이 이번 출시의 핵심 조건입니다.
원문 트윗 2개 보기
first, sorry for the messy rollout. second, when we screw up, we try to make it right. third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future. as usual we will start with pro subscribers.
Stanford AI Lab
The most capable models available in the world are using Terminal-Bench-Science to show it (just one week after the benchmark’s release). Congratulations to @StevenDillmann and team for their remarkable impact!
What a week for Terminal-Bench-Science & AI for Science in general! 🥂 Just 1 week after launch, Terminal-Bench-Science 0.1 is now also the #1 featured benchmark on today's GPT-6 Astra release by @OpenAI. GPT-5.6 Sol → 22.4% GPT-6 Astra → 64.6% The +42.2pp jump is out of
📈 Cybercab의 Starlink V5 통합과 승객용 연결성포스트 3
Cybercab에 Starlink V5가 직접 통합돼 차량 운용과 승객 엔터테인먼트를 위한 연결성을 제공하는 구상이 확산됐으며, 4K 라이브 영상 시청과 fleet vehicle 구매 관심 접수가 함께 언급됐습니다.
세부 내용 보기
- Starlink V5를 Cybercab에 직접 넣으면 통신 서비스가 약하거나 혼잡한 지역에서도 차량 운용과 승객용 엔터테인먼트를 지원한다는 설명이 나왔습니다. Elon Musk는 탑승 중 4K 라이브 영상을 볼 수 있다고 덧붙여 연결성의 용도를 차량 제어와 미디어 소비로 넓혔습니다.
- Tesla가 Cybercab fleet vehicle purchasing에 관심 있는 사람의 정보를 받는 양식을 웹사이트에 게시했다는 소식도 이어졌습니다. 차량 하드웨어와 통신 인프라의 결합이 실제 구매 관심 접수 단계와 함께 언급된 점이 특징입니다.
- Cybercab에는 22인치 터치스크린을 통한 극장·음악·게임 앱과 작업·휴식 공간이 제시됐고, Starlink V5가 이를 위한 통신 기반으로 추가됩니다. 차량 이동 중 콘텐츠 이용을 안정화하려면 셀룰러 상태가 나쁜 장소에서도 연결을 유지해야 한다는 구조입니다.
원문 트윗 2개 보기

Elon Musk
Starlink will allow you to watch 4k live video while riding in Cybercab
Starlink V5 will soon be integrated directly into Cybercab, delivering seamless connectivity for vehicle operations and passenger entertainment even in locations with poor or high traffic cellular service x.com/Tesla/status/2…
Yun-Ta Tsai
The future is here.
The future of transport is safer, more enjoyable & gives you more time back With theater, music & gaming apps on a 22" touchscreen, Cybercab offers you a space to be productive or unwind Soon to be equipped with V5 Starlink hardware, too, for complete connectivity
➖ 오픈 모델의 코드·데이터·레시피 공개포스트 3
K2 Horizon은 가중치뿐 아니라 코드, 학습 데이터, 레시피와 체크포인트를 공개했고, 금융 AI 영역에서도 Ling-3.0-flash-Fin과 FinFIRST의 공개가 접근성과 검증 가능성의 사례로 언급됐습니다.
세부 내용 보기
- 오픈 모델을 둘러싼 기준이 가중치 공개에서 학습 과정 전체의 공개로 이동하고 있습니다. 한 게시물은 코드·평가·로그·가중치까지 공개해야 한다고 요청했고, K2 Horizon은 0.9B부터 375B까지 여섯 모델에 코드, 학습 데이터, 레시피, 체크포인트를 포함했다고 밝혔습니다.
- K2 Horizon은 0.9B, 3.7B, 7B 모델에서 SOTA라고 소개됐으며 휴대전화·노트북·기업 환경을 겨냥했습니다. 모델 파일만 내려받는 방식이 아니라 학습 입력과 재현 절차까지 함께 제공하는 구성이 핵심입니다.
- vLLM은 Ling-3.0-flash-Fin과 FinFIRST의 오픈소스를 금융 AI에 더 접근하기 쉽고 검증 가능하게 만드는 사례로 평가했습니다. 금융용 모델과 전문가 제작 벤치마크를 함께 공개하면 모델 성능을 특정 업무 기준으로 확인할 수 있다는 방향입니다.
코드·학습 데이터·레시피·평가·가중치를 함께 공개하면 모델의 재현성과 검증 가능성이 높아지고, 금융 AI처럼 도메인별 활용을 확장하기 쉬워집니다.
원문 트윗 2개 보기
Santiago
I hope every open model provider out there takes a page from these guys and starts opening up their training code, evals, logs, and weights. Open-weight models are awesome. Open-source models are even better.
When you consider a house, do you just want to be a pay-to-stay tenant like in a hotel, or buy it but without blueprint and document, or own the house alone with all its wiring, plumbing, design and construction details? What if this house is AI ? Today’s release marks the x.com/ifm_ai/status/…
Md Ismail Šojal 🕷️
K2 Horizon just dropped six fully open models from 0.9B to 375B with code, training data, recipes, and checkpoints included. SOTA at 0.9B, 3.7B, and 7B Built for phones, laptops, and enterprise plus full training transparency. The whole fleet is fully open not just the weights. Rare combination. - http:// huggingface.co/collections/IF M/k2-horizon …
➖ TorchSpec와 vLLM의 speculative decoding 학습·서빙 연결포스트 4
TorchSpec가 PyTorch 기반 speculative decoding 초안 모델 학습 프레임워크로 공개됐고, Kimi K3 Draft Collection은 TorchSpec·vLLM·NVIDIA GB200 조합으로 세 초안 모델과 데이터 레시피를 제공했습니다.
세부 내용 보기
- 큰 모델의 모든 생성을 직접 수행하는 대신 초안 모델이 후보 토큰을 만들고 본 모델이 검증하는 speculative decoding을 효율화하려는 흐름입니다. TorchSpec는 이 초안 모델을 PyTorch 환경에서 학습하도록 구성하고, vLLM은 학습 결과를 실제 서빙으로 연결하는 역할을 맡았습니다.
- Kimi K3 Draft Collection에는 EAGLE-3, DFlash2, DSpark 세 초안 모델이 포함됐으며 TorchSpec와 vLLM을 사용해 NVIDIA GB200에서 학습했습니다. 모델 이름뿐 아니라 데이터 레시피도 함께 공개돼 학습 설정을 재현할 수 있는 입력이 추가됐습니다.
- vLLM 게시물은 학습과 서빙의 협업을 강조했고, 별도 사례에서는 Qwen3.6-27B가 M5 Max에서 105 tok/s, 2× MTPLX, 3.5× llama.cpp 수치로 제시됐습니다. 하드웨어별 초안 모델과 실행 경로를 함께 조정하는 방식이 로컬·서버 추론 속도를 좌우하는 구조입니다.
원문 트윗 2개 보기
PyTorch
TorchSpec is a PyTorch-native framework for training speculative decoding draft models. This release, a collaboration with @vllm_project , shows it in action on Kimi K3 @lightseekorg
Releasing the @Kimi_Moonshot K3 Draft Collection — 3 draft models (EAGLE-3, DFlash2, DSpark) trained with TorchSpec and @vllm_project on @NVIDIAAI GB200. 🚀 We also shared the data recipes. More on the blog → https:// lightseek.org/blog/kimi-k3-d raft-collection.html …
vLLM
Love seeing @lightseekorg use TorchSpec + vLLM to train three different K3 draft models, with the recipes open sourced too. Exactly the kind of training-serving collaboration we love to see 🚀
Releasing the @Kimi_Moonshot K3 Draft Collection — 3 draft models (EAGLE-3, DFlash2, DSpark) trained with TorchSpec and @vllm_project on @NVIDIAAI GB200. 🚀 We also shared the data recipes. More on the blog → https:// lightseek.org/blog/kimi-k3-d raft-collection.html …
➖ Qwen3.8-Flash 기반 QwenWork의 자동 deck 생성포스트 2
QwenWork에 Qwen3.8-Flash가 출시됐으며, 파일과 아이디어를 입력하면 논리를 정리하고 deck을 생성하는 작업 흐름과 Standard mode의 1,000 credits당 약 100개 생성 수치가 공개됐습니다.
세부 내용 보기
- QwenWork는 사용자가 자료와 생각을 넣으면 Qwen3.8-Flash가 논리를 자동으로 정리하고 구조화된 deck을 생성하는 방식입니다. 입력 자료의 정리, 내용 구조화, 전문적인 형식의 출력이 한 서비스 안에서 이어져 사용자는 표현과 사고에 집중하도록 설계됐습니다.
- Qwen3.8-Flash는 기존보다 두 배 빠르다고 소개됐고, Standard mode에서는 1,000 credits로 약 100개 deck을 만들 수 있어 개당 약 10 credits가 소요됩니다. 속도와 credit 단가가 함께 제시돼 팀 단위 사용량을 가늠할 수 있는 기준이 마련됐습니다.
- QwenWork 서비스와 함께 Model Studio, Qwen Cloud API 접근 경로가 안내됐습니다. 따라서 자동 deck 생성 기능은 단일 사용자 화면뿐 아니라 팀 사용과 API 연동을 염두에 둔 배포 형태입니다.
원문 트윗 2개 보기
Alibaba Cloud
Qwen3.8-Flash is now live on @qwenwork — twice as fast, and it gets you. Drop in your thoughts and files, and it automatically organizes your logic and generates decks with clear structure and professional polish, so you can focus on thinking and expressing. On Standard mode, 1,000 credits generate about 100 decks — roughly 10 credits each. That's real work done, on a real budget. 👉 Get your team on board and try it together: http:// qwenwork.ai More API access: • Model Studio: https:// click.alibabacloud.com/m/20000002820/ • Qwen Cloud: https:// click.qwencloud.com/m/20000002828/ #QwenWork #AI #Qwen38Flash #AIAgent
Alibaba Cloud
Qwen3.8-Flash is now live on QwenWork — twice as fast, and it gets you. Drop in your thoughts and files, and it automatically organizes your logic and generates decks with clear structure and professional polish, so you can focus on thinking and expressing. On Standard mode, 1,000 credits generate about 100 decks — roughly 10 credits each. That's real work done, on a real budget. 👉 Get your team on board and try it together: http:// qwenwork.ai More API access: • Model Studio: https:// click.alibabacloud.com/m/20000002820/ • Qwen Cloud: https:// click.qwencloud.com/m/20000002828/ #QwenWork #AI #Qwen38Flash #AIAgent
➖ Claude Managed Agent 환경의 파일 기반 동기화포스트 2
Claude CLI에 ant apply가 추가돼 저장소 파일에 선언한 Managed Agent 환경, agents, skills, memory stores와 deployments를 API 리소스와 동기화할 수 있게 됐습니다.
세부 내용 보기
- 에이전트 환경을 콘솔에서 개별 설정하면 저장소와 실행 환경 사이에 차이가 생기기 쉽습니다. ant apply는 저장소 파일에 Claude Managed Agent 환경과 agents, skills, memory stores, deployments를 선언하고 API의 리소스 상태와 맞추는 명령으로 이 문제를 처리합니다.
- 사용자는 설정 파일을 저장소에 두고 ant apply를 실행해 선언된 리소스를 API에 반영합니다. 입력은 저장소의 리소스 정의이고, 처리 단계는 CLI를 통한 동기화이며, 출력은 Claude API에 맞춰진 에이전트 환경과 배포 리소스입니다.
- Modal은 Cursor Cloud Agents에 Modal Sandbox를 연결해 작업별 실행 공간을 제공한다고 밝혔습니다. 두 게시물 모두 에이전트 자체보다 에이전트가 사용할 환경·메모리·배포·샌드박스를 코드와 인프라 단위로 관리하는 방향에 초점을 둡니다.
원문 트윗 2개 보기

ClaudeDevs
We've added `ant apply` to the ant CLI. Now you can declare Claude Managed Agent environments, agents, skills, memory stores, and deployments as files in your repository and keep the API's resources in sync with them using `ant apply`. Docs: https:// platform.claude.com/docs/en/cli-sd ks-libraries/cli/apply …
Run your @cursor_ai Cloud Agents on Modal. Give each Cursor-managed Agent a Modal Sandbox (or a million) hand-crafted for the task.
➖ PyTorch 추론 속도와 대규모 학습 복구성 개선포스트 2
PyTorch AOTI와 KV cache를 결합한 HSTU 추론에서 최대 2.38배 속도 향상이 보고됐고, PyTorch 2.14에서는 프로세스 그룹을 유지한 채 장애 rank를 재구성하는 fault tolerance가 추가됩니다.
세부 내용 보기
- Python backend에서 모델을 실행하면 런타임 오버헤드와 대규모 embedding lookup 비용이 남습니다. PyTorch AOTI는 모델을 Torch C++ runtime에서 실행하고, KV cache는 이미 계산한 상태를 재사용해 반복 계산을 줄이는 방식으로 HSTU 추론 경로를 가속합니다.
- NVIDIA의 HSTU 테스트에서 PyTorch AOTI backend는 Triton Inference Server 배포 기준 Python backend보다 1.14배에서 1.28배 빨랐습니다. 이상적인 all-GPU cache-hit 조건에서 AOTI와 KV cache를 함께 사용하면 2.20배에서 2.38배 속도 향상이 보고됐습니다.
- PyTorch 2.14의 fault tolerance는 대규모 작업에서 rank 하나가 실패했을 때 전체 process group을 해체하고 warm state를 버리는 대신 process-group reconfiguration을 수행합니다. 장애가 난 일부 실행 단위만 재구성해 클러스터 전체의 학습 상태를 보존하려는 구조입니다.
원문 트윗 2개 보기
PyTorch
The PyTorch AOTI backend delivered a 1.14x–1.28x speedup over the Python backend in NVIDIA’s HSTU inference tests when deployed with Triton Inference Server. With the PyTorch AOTI backend and KV cache, @nvidia reports a 2.20x–2.38x speedup in an ideal all-GPU cache-hit scenario. The results come from NVIDIA’s recsys-examples repository, a collection of examples demonstrating best practices for training and deploying generative recommenders on NVIDIA GPUs using PyTorch. It includes optimized HSTU and Semantic ID implementations covering training and inference workflows. For HSTU inference, recsys-examples supports PyTorch AOTInductor to execute the model in the Torch C++ runtime. The NVIDIA developer blog also covers nv-embedding-cache, which provides PyTorch-compatible modules for accelerating large-scale embedding lookups. 🔗 Read the full post:
In PyTorch 2.14, fault tolerance becomes a first-class c10d concept, with in-place process-group reconfiguration. When a rank fails in a large job, the usual recovery is to tear down the process group and restart, which discards warm state across the whole cluster. Backend and
용어 해설
- Terminal-Bench-Science
- — 과학 문제를 다루는 AI 모델의 능력을 측정하는 벤치마크입니다. GPT-6 Astra 출시 직후 평가 항목으로 포함됐으며, GPT-5.6 Sol의 22.4%에서 GPT-6 Astra의 64.6%로 점수가 상승했다는 수치가 인용됐습니다.
- KV 캐시(KV cache)
- — Transformer가 이전 토큰의 Key와 Value를 저장해 다음 토큰 생성 때 다시 계산하지 않도록 하는 메모리 구조입니다. 긴 문맥에서는 전체 캐시를 매번 훑는 비용이 커져 추론 지연과 계산량에 영향을 줍니다.
- 추측 디코딩(Speculative decoding)
- — 작은 초안 모델이 다음 토큰 후보를 먼저 만들고, 큰 모델이 이를 한꺼번에 검증하는 추론 방식입니다. TorchSpec와 Kimi K3 Draft Collection 사례에서는 초안 모델 학습과 서빙을 함께 최적화하는 흐름이 나타났습니다.
- PyTorch AOTI
- — PyTorch 모델을 미리 컴파일해 Python 실행 경로 대신 Torch C++ runtime에서 실행하는 방식입니다. HSTU 추론 테스트에서 Triton Inference Server와 함께 사용됐으며 Python backend 대비 1.14배에서 1.28배 속도 향상이 기록됐습니다.
- 에이전틱 인터넷(agentic Internet)
- — 사람 대신 소프트웨어 에이전트가 웹 서비스와 인프라를 호출하고 작업을 수행하는 인터넷 환경을 뜻합니다. Cloudflare 게시물은 2026년 5월 기준 온라인에서 봇이 사람보다 많아졌다고 밝혔습니다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.