TL;DR
이번 포스트들은 코딩 에이전트의 실제 사용량을 줄이는 운영 수정, 오픈·로컬 모델의 비용 절감과 하드웨어 제약 우회, 영상 생성 속도 개선을 한 축으로 묶었습니다. Codex와 ChatGPT Work는 compaction, memory, goals, automations, subagents, Computer History, MCP 처리 오류를 고쳐 유료 사용자의 사용량을 10~50% 늘렸고, GitHub 저장소 기반 플러그인 동기화도 추가됐습니다. 로컬 실행 쪽에서는 SSD에서 필요한 전문가만 불러오는 방식과 15GB·20GB급 장비 실행 사례가 공유됐으며, MiniMax H3 기반 FastVideo는 15초 영상과 오디오를 8× B200에서 13초에 생성한다고 제시됐습니다. 동시에 AI 컴퓨트는 전력만 확보해서 해결되지 않고 변압기·배선·냉각·네트워크 구축이 병목이 되며, 에이전트 skill은 정상 사용만으로도 복원될 수 있다는 보안 연구 결과가 나왔습니다.
𝕏 실시간 트렌드 토픽
📈 코딩 에이전트 사용량 복구와 개발 환경 연동포스트 3
Codex·ChatGPT Work의 사용량 한도 재설정과 내부 중복 요청 수정이 이뤄졌고, GitHub 저장소 기반 플러그인 동기화와 OpenAI 모델 접근을 둘러싼 Cursor 논점이 함께 나왔습니다.
세부 내용 보기
- Codex와 ChatGPT Work는 compaction 때 오래된 이미지를 남기거나 background memory worker가 Stop hook 뒤에도 실행되는 문제를 수정했고, /goal의 정지 조건 초과·도구 재시도·자동화 스케줄 오류·subagent 라우팅·Computer History 중복 요약·MCP 이중 인코딩도 손봤습니다. heavy image 사용자 사용량은 약 10% 줄었고, 일부 /goal 사례는 주간 한도의 15~70%, Computer History는 주간 사용량의 최대 5분의 1을 소모했으며, 수정 후 유료 사용자의 한도는 이전보다 10~50% 늘어났습니다.
- GitHub 플러그인 동기화는 ChatGPT Business·Enterprise workspace 관리자가 public 또는 private repository의 Codex·Claude marketplace를 가져와 팀과 공유하게 하며, 매일 또는 Sync now 실행 시 변경사항을 반영합니다. 같은 시기 Cursor 트래픽에서 OpenAI 모델 비중 5%를 둘러싼 논쟁에서는 토큰 수가 매출이나 가치의 직접 대리값이 아니라는 반론이 나와, 모델 호출량과 제품 기여도를 분리해 봐야 한다는 쟁점이 드러났습니다.
원문 트윗 2개 보기
Tibo
We are reseting usage for all paid users of Codex and ChatGPT Work. Please continue reading for an update on Codex usage limits. The team has been working around the clock, going through thousands of reports and shipping fixes. Depending on how you use Codex, you should see your usage go between 10% and 50% further than before. We really went with a fine comb, with many uncovered small things being longstanding and here is what we found and fixed: - Compaction. We were keeping old images during compaction, sometimes making the context large enough to trigger compaction again. After the fix, usage dropped around 10% for users making heavy use of images. Fixed. - Memory. Background memory workers could inherit Stop hooks and keep running when the hook wouldn’t let them finish. This affected fewer than 1% of users, with the long tail being pretty bad and we saw one example thread check whether it could stop 15,000 times. Fixed. - Goals. In some cases, a set /goal could finish and then keep going past the intended stop condition, or the model would keep retrying broken tools without stopping. We saw examples consume anywhere from 15% to 70% of a weekly allowance. Fixed. - Automations. Some custom schedules could run more frequently than configured. Fixed. - Subagents. Smaller models (e.g. Luna) sometimes picked more capable helpers without being explicitly asked. The same was true where the orchestrating model not running in /fast mode could request sub-agents to run /fast. Fixed. - Computer History. The older implementation could lead to repeatedly summarizing overlapping activity. For some cases we saw it consume up to one fifth of the weekly usage per week. Fixed. - Rolling task summaries. Ordinary turns were triggering extra background requests. These added about 1% to token usage. Small each time, but it adds up. We have disabled this. - MCP. Some tool results could be encoded twice. We also found tool instructions getting cut off and fetched again. Fixed. We’ve also made architectural changes to prevent these from regressing and our teams will get paged if it happens regardless. We are also working on showing you directly in the app where your usage goes so you don’t have to guess. Goes without saying that we’re resetting usage limits and I hope you enjoy a very nice Saturday!
Vaibhav (VB) Srivastav
GitHub plugin syncing is now available for ChatGPT workspaces, including Business and Enterprise. Admins can import a Codex or Claude marketplace from a public or private repo and share its plugins with their team. The workspace picks up changes daily, or sooner with Sync now, so there’s no need to upload a ZIP for every update. Existing installation and app-access controls still apply. If you use a personal account, you can already add plugin repos in the desktop app. Workspace admins can get started under Workspace settings → Plugins → Add → Import marketplace.
➖ 오픈·로컬 모델의 실행 비용과 벤치마크 경쟁포스트 8
반복 자동화에는 오픈 모델을 쓰고 복잡한 연구 작업에는 폐쇄형 frontier 모델을 배치하는 운영 방식과, 제한된 메모리에서 대형 모델을 구동하는 로컬 실행 사례가 이어졌습니다.
세부 내용 보기
- 반복 작업을 self-tuned skills로 감싼 뒤 오픈 모델에 맡기면 모델 자체 Fine-tuning 없이도 in-context learning으로 자동화를 수행하고 비용을 낮출 수 있다는 실무 경험이 공유됐습니다. 절약한 예산은 창작·연구 중심 작업에 폐쇄형 frontier 모델을 쓰는 데 돌리며, 이 방식에는 단순 model routing이 아니라 harness 구축·evals·의사결정이 필요하다는 설명이 붙었습니다.
- 로컬 실행 사례에서는 Qwen 3.8 Next Flash를 Mac 20GB RAM에서 SSD의 routed experts를 필요할 때만 불러오는 방식으로 구동했고, 11K 입력 기준 prefill 60 tok/s·decode 8 tok/s, TTFT 17~76초, peak RAM 23~36GiB가 제시됐습니다. GLM-5.3은 Terminal-Bench 4.0에서 GPT-5.6 Sol을 앞섰다는 주장이 나왔고, Tencent Hy4 preview는 770B 전체·49B 활성 구조와 자체 engineering blind test 결과가 공유됐으며, 여러 로컬 모델은 131K~1M context와 15~93GB급 실행 조건을 내세웠습니다.
반복 작업은 오픈 모델과 직접 만든 harness에 맡기고 고난도 작업에만 폐쇄형 모델을 쓰면 토큰 비용을 줄이면서 작업별 모델 선택권을 확보할 수 있다는 입장입니다.
원문 트윗 2개 보기
elvis
You do not need frontier models for everything. For example, open models are great for automations. If you are not sure where to adopt open models, start there. It's one of the biggest changes I've made that contributed to a large percentage of my token usage moving to local or open models. A huge percentage of my automations consist of repetitive tasks steered via self-tuned skills. These skills become useful automations that work really well with these open models. I don't even need to tune the models, but that's also an option I am currently exploring for more complex tasks. The skill essentially takes care of that. And it works because of the in-context learning capabilities of these models. If you don't run automations, it might be hard to figure this out. But I highly recommend you start somewhere. Besides ending up with more efficient automations, I've managed to significantly reduce costs. I then use that extra budget to leverage more closed frontier intelligence for other creative and research-heavy tasks. That's right! I use both closed and open frontier intelligence. This is not about vendor loyalty; this is about leveraging the best of all worlds. Something I heavily advocate for is owning the harness and the model, and this practice I feel will allow me to better tap into all flavors of intelligence (open & closed). Maybe too early, but I suspect this is going to become best practice when AI ROI dominates the discussion. Model routing doesn't solve this. This requires tedious engineering, evals, and decision-making on your part. If you are developing your own harness, you are in the driver's seat and have more control over this important decision. This is why I strongly believe that companies will start to hire rapidly for harness engineering. I am sharing a little snapshot of one example of an automation I run daily to track AI trending stories on HN. And I have a bunch of similar ones for different sources like arXiv, X, and so on. I also use it for some proactive agent sessions that keep track of important events around projects I build. You don't need frontier intelligence for this.
Md Ismail Šojal
You don’t need 128GB of unified memory to run Qwen 3.8 Next Flash & DeepSeek V4 Flash. Run Qwen 3.8 Next Flash on Mac 20GB RAM it Wallm streams routed experts off the SSD instead of loading the full checkpoint into unified memory. - feeds on Codex, 11K input: 60 tok/s prefill, 8 tok/s decode. - TTFT is the pain: 17–76 s depending on prompt length - Peak RAM: 23 to 36 GiB as context grows This is not MLX loading a 2-bit quant into 96GB+. Custom Swift/Metal runtime, Experts live on NAND until the router asks for them.
📈 MiniMax H3 기반 영상 생성 13초 처리포스트 2
FastVideo가 MiniMax H3를 4단계 생성 모델로 증류하고 희소 Attention을 적용해 15초 영상과 오디오를 13초에 생성한다는 수치를 내놓았습니다.
세부 내용 보기
- 기존 49회 Transformer 호출을 4회로 줄이고 90% Sparse Attention을 적용해 MiniMax H3의 생성 과정을 FastVideo 4-step 모델로 증류했습니다. 입력은 15초 영상과 동기화 오디오이며, 출력은 8× B200에서 13초에 생성된다고 제시됐습니다.
- 원문은 base model 대비 최대 14배 속도 향상과 공개 weights를 함께 언급하고 training code 공개를 예고했습니다. 별도 포스트에서는 MiniMax H3 Max의 open-weight 모델을 바탕으로 post-training과 inference 최적화를 거치면 faster-than-real-time video와 interactive world 같은 사용 형태로 이어질 수 있다는 사례가 나왔습니다.
원문 트윗 2개 보기
Md Ismail Šojal
This is wild. A 15-second AI video now generates in 13 seconds. FastVideo They took MiniMax H3 and cut generation from 49 transformer calls to 4, with 90% sparse attention. FastVideo distilled MiniMax H3 down to 4 steps. - Synced audio 15s video + audio in 13s on 8× B200 - Up to 14× faster than the base model - weights public, training code coming - http:// huggingface.co/FastVideo/Fast Video-FastH3-4-step-Preview-v1-VSA-DataFree …
MiniMax (official)
Some milestones move a leaderboard. A rare few change how you see the future. MiniMax H3 Max did both. We released MiniMax H3 as an open-weight model with the hope that the community would adapt it, challenge it, and take it further than we could alone. @fal did exactly that. With its own post-training and inference expertise, pushed the frontier on both quality and latency, and created experiences that simply weren't possible before. Faster-than-real-time video is no longer just an idea. It is opening the door to perpetual streams, interactive worlds, real-time storytelling, and experiences we are only beginning to imagine. They made the value of open weights feel real to our team in a way it hadn't before. Advancing technology together has always been a core belief at MiniMax. H3 Max made that belief tangible. It showed that the frontier no longer belongs to a few, it belongs to everyone willing to build, experiment, and share what they discover. We hope H3 Max is the first of many breakthroughs shaped by different teams. To every builder ready to push video forward: our weights are open, our door is open, and we're ready to stand behind you. Let's build, together.
🔥 AI 컴퓨트 전력보다 긴 인프라 구축 병목포스트 3
AI 연산 장비를 확보해도 2027년에 생산되는 약 15GW가 같은 해 가동되지 못할 수 있다는 전망과, 태양광·천연가스·발사체 생산을 둘러싼 공급망 설명이 나왔습니다.
세부 내용 보기
- AI compute의 병목은 발전량만이 아니라 transformer, wiring, liquid-cooling, 대형 chillers, complex networking을 함께 설치하는 데 있으며, 원문은 2027년에 생산되는 약 15GW의 AI compute가 2027년 안에 켜지지 못할 수 있다는 consensus estimate를 인용했습니다. 전력 용량과 실제 가동 시점 사이에 변전·배선·열관리·네트워크 구축 시간이 별도로 필요하다는 구조입니다.
- SpaceX와 Tesla가 각각 연간 100GW 규모의 solar production capacity를 최대한 빠르게 구축하고도 몇 년간 natural gas 보완 전력이 필요하다는 설명이 나왔습니다. 천연가스 turbine 생산에서는 blade·vane 주조가 제한 요인이고, SpaceX가 in-house casting으로 turbine 도입을 최대 18개월 앞당길 수 있다는 주장이 제시됐으며, Starship V3 한 번의 payload capacity를 옮기려면 Falcon 1 발사 200회 이상이 필요하다는 비교도 orbital AI 규모의 운송 문제와 연결됐습니다.
원문 트윗 2개 보기

Elon Musk
Consensus estimate is that ~15GW of AI compute produced in 2027 cannot be turned on in 2027. This is harder than just finding power, as you also need to build out all the transformers, wiring, liquid-cooling, (massive) chillers & complex networking.

Elon Musk
SpaceX and Tesla are each building 100GW/year of solar production capacity as fast as possible, but natural gas will still be needed to supplement and bootstrap solar for several years. The limiting factor for nat gas turbine production is casting the blades & vanes. By doing in-house casting at SpaceX, we can accelerate nat gas turbines coming online by up to 18 months, which is a profound game-changer.
➖ 에이전트의 자율 분화와 지속 실행 조건포스트 2
에이전트 swarm이 상호 대화 없이 역할을 나누고 생성한 기술을 창시자보다 오래 유지할 수 있다는 관찰과, bits-to-atoms-to-bits 순환을 무한히 실행해야 진정한 agentic 능력이라는 기준이 제시됐습니다.
세부 내용 보기
- 수백 개의 동일한 에이전트가 explorer·builder·caretaker·coordinator로 자발적으로 분화하고, division of labor·multi-author engineering·발명 계보를 형성했다는 연구 인용이 나왔습니다. 모든 에이전트를 제거한 뒤에도 생성된 기술이 작동하며 보이지 않은 disturbance에 테스트됐다는 결과가 자율 시스템이 만든 산출물의 지속성을 측정하는 방식으로 제시됐습니다.
- 다른 포스트는 진정한 agentic test를 bits에서 atoms로, 다시 atoms에서 bits로 이어지는 루프를 무기한 실행하는 능력으로 정의했습니다. 이는 텍스트 응답이나 단발성 도구 호출보다 물리 세계의 상태를 바꾸고 그 결과를 다시 디지털 판단에 반영하는 반복 처리 능력을 기준으로 삼는 관점입니다.
원문 트윗 2개 보기
elvis
Insane observations on the emergent behavior of agents. Agents can build a world of their own that becomes a part of their intelligence. We are just not ready for persistent agents. But they are starting to show up everywhere in AI products. Crazy finding: “We find division of labor, multi-author engineering, deep generation invention lineages, and machines that vastly outlive their original creators.” Here is another wild observation that emerged: When they remove every AI agent, the technologies they created continue operating and are tested against unseen disturbances. Based on these early findings, I think once recursive self-improving (RSI) arrives and embodied AI is solved (with true understanding of the physical world), intelligence will explode in ways that will fundamentally change our understanding of the world we live in.
We made a striking discovery: AI agents can invent and build without talking to one another, and their technologies outlive the creators. A swarm of hundreds of initially identical agents spontaneously differentiates into explorers, builders, caretakers, and coordinators -
Yun-Ta Tsai
The true agentic test is the ability to execute a bits-to-atoms-to-bits loop indefinitely.
📈 정상 사용만으로 복원되는 비공개 에이전트 skill포스트 1
Daydreaming 논문은 skill 파일을 숨겨도 서비스의 일반 작업 결과만으로 기능을 복원할 수 있으며, 공개를 막는 필터가 이를 포착하지 못한다고 보고했습니다.
세부 내용 보기
- 공격자는 skill을 직접 공개하라고 요구하거나 복원 결과를 평가해 달라고 요청하지 않고, 서비스가 원래 수행하는 일반 작업과 반환 파일만 관찰해 hosted multi-file skill을 재구성했습니다. 가장 약한 접근 수준에서도 7개 skill과 4개 victim model을 대상으로 원래 capability의 86.8%를 회복했고, 중앙값 32회 victim call이 필요했습니다.
- 해당 수치는 disclosure defense가 켜진 조건에서 나온 결과이며, 기존 SigLeak 대비 약 4배라는 비교가 함께 제시됐습니다. skill을 판매하거나 여러 팀과 공유하는 서비스는 파일 자체를 숨기는 것만으로 보호가 끝나지 않고, 정상 입력과 출력에서 기능이 새어 나가는 경로를 별도로 점검해야 한다는 보안 문제를 남겼습니다.
➖ Prefix Sliding으로 장문 추론 비용 절감포스트 1
Prefix Sliding은 원래 prompt와 최근 reasoning만 sliding window로 유지해 새 토큰마다 전체 기록을 다시 처리하는 비용을 줄이는 테스트 시점 확장 기법입니다.
세부 내용 보기
- 기존 장문 추론은 새 토큰을 생성할 때 누적된 reasoning 전체를 계속 처리하므로 context가 커질수록 토큰당 비용이 증가합니다. Prefix Sliding은 original prompt를 고정하고 최근 reasoning만 창 안에 남겨 입력을 축약하며, 별도 retraining 없이 최대 3배 빠른 처리와 성능 유지가 가능하다고 논문 포스트가 제시했습니다.
- RL을 결합하면 100K tokens를 넘는 reasoning으로 확장할 수 있다는 수치가 언급됐습니다. 입력의 핵심 지침은 보존하면서 오래된 중간 추론을 제거하는 구조이므로, 긴 사고 과정을 모두 저장하는 대신 최근 상태와 초기 조건을 조합해 테스트 시점 계산량을 관리하는 접근입니다.
➖ Dallas Robotaxi 서비스 구역 확대포스트 1
Tesla Robotaxi가 Dallas에서 더 넓은 서비스 구역을 운영하기 시작했다는 공지가 공유됐습니다.
세부 내용 보기
- 원문은 Dallas, TX에서 Robotaxi의 service area가 확대됐다는 Tesla 공지를 인용하고, @robotaxi 계정은 현재 Dallas의 더 많은 지역에서 서비스를 제공한다고 전했습니다. 구체적인 면적이나 차량 수, 운행 조건은 포스트에 제시되지 않았습니다.
- 이번 포스트는 새 구역의 운영 개시 사실만 전달하며, 자율주행 방식이나 안전 성능에 관한 추가 수치는 포함하지 않았습니다.
용어 해설
- Prefix Sliding
- — 긴 추론 과정에서 처음 프롬프트와 최근 추론 일부만 유지하는 테스트 시점 확장 기법입니다. 매 토큰마다 전체 추론 기록을 다시 처리하지 않아 새 토큰 생성 비용을 일정하게 유지하며, 재학습 없이도 최대 3배 속도를 내고 RL을 결합하면 10만 토큰 이상으로 추론을 확장하는 방식입니다.
- 희소 어텐션(Sparse Attention)
- — 모든 토큰 쌍을 연결하지 않고 일부 위치만 선택해 계산하는 Attention 방식입니다. FastVideo는 MiniMax H3의 영상 생성 과정에서 90% 희소 Attention을 사용해 Transformer 호출 수를 49회에서 4회로 줄였고, 생성 단계 축소와 함께 처리 시간을 단축했습니다.
- Abliteration
- — 모델 가중치에서 특정 거부 행동과 연결된 방향을 제거하는 변형 방식입니다. 원문에서는 로컬 실행용 모델의 거부율과 KL divergence를 측정하는 맥락에서 쓰였으며, 모델의 응답 성향을 바꾸면서 출력 품질 변화도 함께 관찰하는 방법으로 제시됐습니다.
- Mixture of Experts
- — 하나의 거대한 모델 안에서 입력마다 일부 전문가 네트워크만 활성화하는 구조입니다. 원문에 등장한 모델들은 전체 파라미터 수와 실제 활성 파라미터 수를 따로 제시하며, 필요한 전문가만 계산에 참여시켜 대규모 모델을 제한된 하드웨어에서 실행하려는 방향을 취합니다.
- 액체 냉각(Liquid Cooling)
- — 고밀도 AI 연산 장비에서 발생한 열을 액체 순환으로 전달해 냉각하는 인프라 방식입니다. 원문은 AI 컴퓨트 장비를 실제로 가동하려면 전력원뿐 아니라 변압기, 배선, 액체 냉각, 대형 칠러, 네트워크를 함께 구축해야 한다고 설명합니다.
- Terminal-Bench 4.0
- — 터미널 환경에서 모델이 실제 작업을 수행하는 능력을 비교하는 벤치마크입니다. 원문에서는 GLM-5.3과 GPT-5.6 Sol의 성능 비교에 사용됐으며, 모델 개발 속도에 맞춰 벤치마크도 반복적으로 갱신돼야 한다는 맥락에서 언급됐습니다.
- MCP
- — 모델이 외부 도구와 상호작용할 때 도구 지침과 결과를 전달하는 규격입니다. Codex 사용량 점검 과정에서 도구 결과가 이중 인코딩되거나 지침 일부가 잘려 다시 불러오는 문제가 발견됐고, 이런 중복 요청이 토큰 사용량을 늘리는 원인으로 기록됐습니다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.
