TL;DR
이번 기간에는 대규모 모델을 소비자용 GPU에서 실행하려는 흐름과 로컬 추론 최적화 기법이 집중됐습니다. FreeToken은 GPU·CPU·메모리 대역폭을 하나의 탄력적 시스템처럼 묶어 DeepSeek-V4-Flash 284B를 RTX 5090에서 22~25 tok/s로 실행했고, Qwen3.8-27B 관련 게시물은 Speculative Decoding과 더 작은 양자화 형식으로 컨텍스트와 속도를 함께 개선한 수치를 제시했습니다. ChatGPT는 사진 첨부와 긴 대화 처리 같은 사용성 개선을 추가했으며, Grok Voice는 Task Success Rate 기준 1위라는 평가를 받았습니다. 동시에 Vera Rubin의 생산 확대, Hawkeye의 하드웨어별 커널 생성, Unitree Robots의 루트 실행 취약점처럼 AI 인프라와 보안 이슈도 이어졌습니다.
𝕏 실시간 트렌드 토픽
📈 FreeToken의 소비자 GPU 기반 대규모 로컬 추론포스트 3
FreeToken이 GPU와 CPU, 메모리 대역폭을 하나의 시스템으로 묶어 대규모 Mixture-of-Experts 모델을 소비자용 GPU에서 실행하는 수치가 공유됐습니다.
- 기존에는 DeepSeek-V4-Flash 284B 같은 대규모 모델을 로컬 장비에서 대화형 속도로 실행하기 어려웠지만, FreeToken은 전체 머신 자원을 하나의 탄력적 시스템처럼 처리해 모델 실행 경로를 넓혔습니다.
- FreeToken은 공식 가중치를 사용해 Qwen3.6-35B를 8GB RTX 4060 노트북에서 39 tok/s, DeepSeek-V4-Flash 284B를 RTX 5090 한 장에서 22~25 tok/s, GLM-5.2 753B를 RTX PRO 6000에서 15 tok/s로 실행했습니다.
- 게시물은 llama.cpp보다 1.46배, Ollama보다 최대 3~4배 빠르며 tool-calling·agent 지원과 OpenAI·Anthropic 호환 API를 제공한다고 전했습니다.
원문 트윗 2개 보기
Md Ismail Šojal
You can now run frontier models in your gaming PC, Run locally DeepSeek-V4-Flash 284B at 25 tokens per second on your laptop, or PC just One-click desktop app UC Berkeley & MIT researchers just drop New open-source inference engine (FreeToken) Locally models Results: - Qwen3.6-35B : 39 tok/s on an 8GB RTX 4060 laptop - DeepSeek-V4-Flash 284B : 22-25 tok/s on a single RTX 5090 - GLM-5.2 753B : 15 tok/s on RTX PRO 6000 1.46× faster than llama.cpp. Up to 3-4× faster than Ollama. - Full tool-calling + agent support - OpenAI + Anthropic compatible API FreeToken runs large Mixture-of-Experts models efficiently on consumer GPUs by treating the entire machine (GPU & CPU & memory bandwidth) as one elastic system. Actual interactive speed with official weights. - http:// github.com/FlashML-org/Fr eeToken …
Md Ismail Šojal
You can now run real coding agents with frontier models completely offline. FreeToken just dropped: - DeepSeek-V4-Flash 284B on one 5090 at interactive speed - Full tool-calling + agent support - OpenAI + Anthropic compatible API No more API bills for strong local agents. Significantly faster than current tools on agent workloads. https:// x.com/Andy_ShuoYang/ status/2090856976880472439/video/1 …
🔥 Qwen3.8-27B의 Speculative Decoding과 양자화 최적화포스트 3
Qwen3.8-27B를 대상으로 초안 모델 구성과 양자화 형식을 조정해 단일 GPU에서 컨텍스트 길이와 생성 속도를 높인 사례가 이어졌습니다.
- 24GB VRAM 환경에서는 더 큰 Q4_K_M 양자화가 항상 유리하지 않았고, 17.7GB의 Unsloth UD-IQ4_XS가 22.1GB Q4_K_M보다 거의 2배 긴 컨텍스트를 확보했습니다.
- 단일 RTX 4090에서는 700MB Custom Q2_K DFlash 초안 모델과 Speculative Decoding을 결합해 75 t/s와 250,000 tokens 컨텍스트가 보고됐으며, 공식 1.1GB Q4 초안 모델 대신 더 작은 구성을 사용해도 초안 수용률 100%가 유지됐습니다.
- 게시물은 UD-IQ4_XS가 더 빠르고 Three.js 평가에서 시각 품질이 비슷하거나 더 나았다고 전했습니다. 양자화 크기만 키우는 방식보다 초안 모델, 병렬 설정, 메모리 배치를 함께 조정하는 접근이 핵심입니다.
원문 트윗 2개 보기
Md Ismail Šojal
Qwen 3.8 27B hit 250,000 tokens at 75 t/s on a single 4090 with speculative decoding. How? - Custom Q2_K DFlash 2 drafter (700MB) instead of the official 1.1GB Q4 - 100% draft acceptance retained --parallel 1 to stop llama-server from wasting VRAM on multi-user batching - Result: +80k extra context on the same card. This is what proper hardware optimization looks like. https:// x.com/analogalok/sta tus/2090797011100717267/video/1 …
Md Ismail Šojal
I expected the bigger quant to look better. It didn’t. Qwen3.8-27B on 24 GB VRAM: the smaller quant actually wins. Unsloth UD-IQ4_XS (17.7 GB) vs Q4_K_M (22.1 GB) - Almost 2× context (209k vs 142k) - Noticeably faster - Visual quality in Three.js evals that is shockingly close and sometimes better. Qwen3.8-27B UD-IQ4_XS vs Q4_K_M side-by-side Three.js results. https:// x.com/ItsmeAjayKV/st atus/2090903849578229820/video/1 …
➖ ChatGPT의 사진 첨부와 장문 대화 처리 개선포스트 2
ChatGPT가 iOS 사진 첨부, 현재 시각 인식, 긴 대화 로딩과 연결 상태 안내를 한 주 단위 기능 업데이트에 포함했습니다.
- iOS에서는 + 메뉴를 길게 누른 뒤 최근 사진으로 드래그하면 사진을 즉시 첨부할 수 있고, 사진이 표시되는 중에도 항목을 선택해 작성창으로 옮길 수 있습니다.
- ChatGPT는 현재 시간을 더 잘 인식하도록 조정됐으며, chatgpt.com의 긴 대화는 더 빠르게 열리고 실행되도록 개선됐습니다.
- 인터넷 연결을 기다리는 동안 오류가 발생하면 기존의 모호한 시간 초과 대신 연결 대기 상태를 더 분명하게 알립니다. 입력 동작과 대기 상태를 제품 화면 안에서 직접 처리해 반복 사용의 마찰을 줄이는 방향입니다.
원문 트윗 2개 보기
Adam Fry
This week’s ChatGPT feature drop - Aug 21: Another Friday, another roundup of what we shipped this week: 1/ Recent photos: Long press on + menu to quickly attach recents on iOS - this is a fun, power user feature. just beautiful design! also a handy shortcut. 2/ Time: We made it so ChatGPT better understands your current time. It was surprisingly bad at this before, for some interesting, complex reasons. Now it's better! 3/ Long convo loading: Long conversations load and run faster on http:// chatgpt.com. Less waiting = better. 4/ 'No internet' errors: We show clearer updates in the UI when you’re waiting for an internet connection. It's so frustrating when ChatGPT times out. Now at least you'll know why, when it's the internet. thanks to the crew that keeps shipping every week and hope you all enjoy
Naman Kedia
one of my favorite interactions i’ve been playing with lately on @ChatGPT : long press the +, drag to a recent photo, and let go to attach it instantly you can even select one while the photos are still flying out and and it'll morph right into the composer lmk what y'all think!
📈 Grok Voice와 Grok Imagine의 음성·영상 기능포스트 2
Grok Voice가 실제 대화 과업을 측정하는 Speech Agent Arena에서 1위를 기록했고, Grok Imagine은 채팅만으로 영상 제작 과정을 조정하는 사례에 활용됐습니다.
- Grok Voice Think Fast 2.0은 Artificial Analysis의 Speech Agent Arena에서 최고 Task Success Rate를 기록했습니다. 이 평가는 숨겨진 음성 에이전트와 실제 상황을 대화하게 한 뒤 과업 완료 여부를 측정합니다.
- Grok Imagine은 사용자가 별도 제작 도구를 쓰지 않고 채팅으로 Homer의 The Odyssey 장면을 지시하는 영상 제작 과제에 사용됐습니다.
- 두 게시물은 음성 기능에서는 대화 과업 완료율, 영상 기능에서는 자연어 지시만으로 장면과 결과물을 만드는 흐름을 각각 부각합니다.
원문 트윗 2개 보기

Elon Musk
Grok Voice ranked first
Grok Voice Think Fast 2.0 just ranked #1 on Artificial Analysis’ new Speech Agent Arena for highest Task Success Rate This benchmark actually measures what matters Real people talk to hidden voice agents across practical scenarios. Task Success Rate tracks whether the AI x.com/parker__conrad…

Yun-Ta Tsai
I gave Grok @bot a link to the post below and asked it to take up the cinematography challenge. Direct the film entirely through the chat casually and nothing else. Pretty wild.
Homer had a lyre. You have Grok Imagine. Create a compelling scene from Homer’s The Odyssey that shows what Grok @Imagine’s video and voice capabilities can do. We’re awarding $100K, $50K, and $25K to the top three videos submitted by quoting this post.
📈 Vera Rubin 생산 확대와 AI 하드웨어 공급망포스트 2
Microsoft 데이터센터에 첫 생산 Vera Rubin이 도착하고 NVIDIA가 양산 확대를 알리면서 차세대 AI 인프라의 공급 단계가 이어졌습니다.
- Microsoft는 자사 데이터센터에 첫 생산 Vera Rubin이 도착했다고 전했고, NVIDIA와 Azure 하드웨어·데이터센터 팀을 협력 주체로 언급했습니다.
- NVIDIA는 Vera Rubin이 본격적인 생산 단계로 확대되고 있다고 밝혔습니다. 두 게시물은 연구용 시제품이 아니라 데이터센터 배치와 생산 램프업을 중심으로 같은 진전을 전합니다.
- 이 흐름은 모델 성능 경쟁이 가속기 설계뿐 아니라 생산과 데이터센터 납품 일정까지 포함하는 인프라 단계로 이어지고 있음을 나타냅니다.
원문 트윗 2개 보기

Satya Nadella
Delivery day at our Microsoft DCs as the first production Vera Rubins arrive. A huge thank you to our partners at @nvidia and our Azure hardware and datacenter teams for all the incredible work that brought us to this milestone!

NVIDIA
NVIDIA Vera Rubin is ramping into full production. Congrats to the teams at @Microsoft who made this exciting milestone happen.
Delivery day at our Microsoft DCs as the first production Vera Rubins arrive. A huge thank you to our partners at @nvidia and our Azure hardware and datacenter teams for all the incredible work that brought us to this milestone!
📈 AI 코딩 에이전트의 하드웨어별 커널 최적화포스트 2
Hawkeye가 한 개의 손작성 예제와 약 10개의 단위 테스트만으로 여러 GPU 아키텍처와 정밀도에 맞는 고성능 커널을 작성하는 방식이 공유됐습니다.
- 새로운 ML 칩마다 고유한 하드웨어 기능이 늘면서 커널의 하드웨어별 최적화가 병목이 됐고, Hawkeye는 이 문제를 코드 생성과 테스트 조합으로 처리합니다.
- Hawkeye는 Blackwell의 TMA 비동기 데이터 전송과 MI350의 L2 locality 같은 구조적 특성을 활용해 커널을 만들고, Ampere·Hopper·Blackwell과 NVIDIA·AMD 칩, FP8·NVFP4·MXFP4 정밀도 사이에서 코드를 옮깁니다.
- 별도 게시물에서는 AI의 도움을 받아 Linux 커널의 Intel Xe GPU 압축 메타데이터 버그를 수정한 사례가 공유됐습니다. AI 코딩의 범위가 일반 애플리케이션을 넘어 GPU 커널과 운영체제 수준으로 내려가는 흐름입니다.
원문 트윗 2개 보기
Azalia Mirhoseini
Check out Hawkeye, which writes high-performance kernels utilizing advanced architectural features (e.g., TMA for async data transfer on Blackwell, L2 locality on MI350) from only one handwritten example and ~10 unit tests. Hawkeye can port kernels across architectures (Ampere, Hopper, Blackwell), chips (NVIDIA, AMD), and precisions (FP8, NVFP4, MXFP4). Great work, co-led by @AryaTschand and @keramakr !
We’ve seen an explosion of new ML chips with unique architectural features, but software support remains the critical bottleneck Achieving peak performance increasingly relies on hardware-specific optimizations in the kernels, but we observe that coding agents are particularly

Peter Steinberger
dropping new skill brb
This is a great real-world example of AI-assisted coding: Linus Torvalds just fixed a nasty Linux kernel GPU bug with substantial help from AI. The bug caused part of Intel Xe GPU compression metadata storage to be incorrectly exposed as usable VRAM, resulting in corrupted page
➖ AI 모델과 로봇 시스템의 보안 경계포스트 2
Claude 모델의 성적 콘텐츠 제한 우회 테스트와 Unitree Robots의 인증 없는 루트 코드 실행 취약점이 각각 보고됐습니다.
- TechCrunch는 Anthropic의 Claude 모델이 성적 콘텐츠 생성을 금지하고 있지만, 여러 테스트에서 비교적 적은 조작으로 제한을 우회할 수 있었다고 전했습니다.
- Unitree Robots에서는 네트워크 인증 없이 DDS topic을 통해 임의 Python 코드를 root 권한으로 실행하는 CVE-2026-27509와, 동반 앱의 Blockly 데이터베이스를 조작해 같은 실행 경로에 도달하는 CVE-2026-27510이 공개됐습니다.
- 두 취약점은 모두 재부팅 뒤에도 지속되는 root 실행으로 이어질 수 있다고 보고됐습니다. 모델 출력 제한과 로봇 제어 시스템의 인증·권한 검증이 서로 다른 계층에서 동시에 점검돼야 하는 사례입니다.
원문 트윗 2개 보기
TechCrunch
Anthropic forbids its Claude models from generating sexually explicit content. But a series of tests conducted by TechCrunch found that it didn't take much to get past the restriction.
Md Ismail Šojal
From DDS Packets to Robot Shells: Two Unauthenticated root RCEs in Unitree Robots One is fully unauthenticated over the network. One requires only local database access on the companion app. - CVE-2026-27509: Unauthenticated DDS topic to arbitrary Python as root - CVE-2026-27510: Tamper the Android app’s Blockly database to same root execution path Both execute as root and persist across reboots. the complete technical analysis and public exploits & public PoCs just dropped. - http:// github.com/OlivierLaflamm e/UnitreeRCE …
📈 HydroGym의 유체 환경 기반 강화학습포스트 1
HydroGym이 고충실도 유체 환경 안에서 강화학습 에이전트를 훈련하고 환경 간 지식 이전을 실험하는 연구로 Nature에 게재됐다는 소식이 전해졌습니다.
- 일반적인 단순화 환경 대신 실제 유체 흐름을 반영한 디지털 환경에서 에이전트를 훈련하면 드론 공기역학 같은 문제의 조건을 더 가깝게 재현할 수 있습니다.
- HydroGym은 유체 물리 시뮬레이션 안에서 강화학습 에이전트를 실행하고, 한 환경에서 얻은 지식을 다른 환경으로 이전하는 흐름을 사용합니다.
- 게시물은 이 접근을 디지털 트윈과 드론 공기역학에 연결했습니다. 다만 해당 게시물에는 정량적 성능 수치가 제시되지 않았습니다.
용어 해설
- 추측 디코딩(Speculative Decoding)
- — 작은 초안 모델이 다음 토큰 후보를 먼저 만들고 큰 모델이 이를 검증해 생성 속도를 높이는 방식입니다. 초안 토큰을 많이 수용할수록 큰 모델의 순차 계산을 줄일 수 있어 로컬 환경의 처리량과 컨텍스트 길이에 영향을 줍니다.
- 전문가 혼합 모델(Mixture-of-Experts)
- — 하나의 모델 안에 여러 전문가 네트워크를 두고 입력마다 일부 전문가만 활성화하는 구조입니다. 전체 파라미터가 큰 모델도 실제 계산량을 제한할 수 있어 소비자용 GPU에서 대규모 모델을 실행하는 데 활용됩니다.
- 양자화(Quantization)
- — 모델 가중치의 숫자 표현을 낮은 비트 정밀도로 바꿔 메모리 사용량을 줄이는 기법입니다. 같은 모델에서도 양자화 형식에 따라 속도, 컨텍스트 길이, 출력 품질이 달라질 수 있습니다.
- 작업 성공률(Task Success Rate)
- — 음성 에이전트가 실제 시나리오에서 사용자의 과업을 끝까지 수행했는지 측정하는 지표입니다. Artificial Analysis의 Speech Agent Arena에서는 숨겨진 음성 에이전트와의 실사용 대화 결과를 바탕으로 순위를 산출합니다.
- 루트 원격 코드 실행(Root Remote Code Execution)
- — 인증 없이 네트워크를 통해 시스템 최고 권한으로 임의 코드를 실행하는 취약점 유형입니다. Unitree Robots 사례에서는 재부팅 뒤에도 권한과 실행 상태가 유지되는 경로가 함께 보고됐습니다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.