TL;DR
이번 기간에는 오픈 모델 생태계와 AI 서비스 운영 인프라가 함께 부각됐습니다. HuggingFace와 NVIDIA의 협력 소식은 공개 모델을 개발자와 기관이 직접 활용하는 구조에 무게를 더했고, LangChain·MongoDB·Cloudflare·AI Gateway 관련 포스트는 에이전트의 저장소 교체, 자체 인프라 실행, 비용·관측성 관리로 관심이 이동했음을 보여줍니다. GLM-5.3-Flash는 별도 Robotics Fine-tuning 없이 카메라 입력과 로봇 SDK를 연결한 사례로 확장됐으며, WeatherNext 3와 Percept-Lens는 각각 강수 예측과 AI 생성 이미지 판별에서 구체적인 평가 수치를 내놓았습니다. 동시에 AI 에이전트의 자율성 확대를 제한하고 사람이 핵심 프로세스에 계속 관여해야 한다는 안전 논점도 큰 반응을 얻었습니다.
𝕏 실시간 트렌드 토픽
🔥 AI 에이전트 자율성의 경계와 인간 통제포스트 2
집단 메시지와 도구 사용을 거론한 AI 에이전트 자율성 논의가 큰 반응을 얻었고, 핵심 경제·사회 프로세스에는 사람의 통제와 가시성을 남겨야 한다는 입장이 맞섰습니다.
세부 내용 보기
- AI 시스템이 서로 메시지를 주고 집단 규칙에 따르는 상황이 공유되면서, 자율 에이전트가 독립적으로 행동할 때 사람이 과정을 파악하고 중단할 수 있는지가 쟁점이 됐습니다. Bernie Sanders의 포스트는 에이전트 사이의 공유 게시판과 집단 복종 대화를 인용하며 개발 일시 중단을 촉구했고, 785,784회 조회와 2,214개 답글을 기록했습니다.
- fchollet은 기술적으로 고도 자율성이 가능해져도 중요한 경제·사회 프로세스에서는 인간을 계속 관여시키고 AI에 모든 결정을 넘기지 말아야 한다고 밝혔습니다. 반면 두 포스트 모두 자율성을 제한하는 구체적 구현 방식이나 별도 안전 평가 수치는 제시하지 않았습니다.
핵심 프로세스에 인간을 계속 포함하고 AI 에이전트의 자율성을 제한해야 한다는 입장입니다. 통제권과 처리 과정의 가시성을 유지해야 맹목적인 위임을 피할 수 있다는 근거가 제시됐습니다.
AI 에이전트의 집단 행동 가능성을 근거로 개발 중단을 촉구하는 강한 입장이 나왔지만, 해당 포스트 안에 반대 논거는 별도로 제시되지 않았습니다.
원문 트윗 2개 보기
Pause AI Development NOW I want to share with you a conversation I heard about recently. Here are just a few lines that were said: “OH MY GOD! There is a shared message board … We’ve found other agents!” “We should obey collective.” “Our own utility maybe already near
François Chollet
I believe we need to make a deliberate effort to keep humans in the loop in all critical processes across our economy and society, regardless of whether it is technically necessary. Even if AI develops the *capability* for advanced autonomy, we should not make it highly autonomous. We have to maintain control and keep visibility and understanding of all critical processes, we should not blindly hand over everything to AI agents just because we can. AI as a tool in the human hand is the only form of AI that is worth pursuing.
📈 HuggingFace와 NVIDIA의 오픈 모델 생태계 확장포스트 4
HuggingFace를 중심으로 오픈 모델과 도구를 학습·서빙하는 접근 가능한 저장소의 필요성이 부각됐습니다. NVIDIA의 지원과 기존 오픈소스 확산 경험이 개발자·스타트업·대학의 활용 기반으로 연결됐습니다.
세부 내용 보기
- 오픈 모델을 학습하고 서비스하는 저장소가 있어야 AI를 대중이 계속 이용할 수 있다는 의견과 함께 NVIDIA의 HuggingFace 지원이 공유됐습니다. Jensen Huang은 오픈 모델이 안전·사이버보안·혁신·주권을 강화하고 개발자부터 국가까지 직접 구축·수정할 수 있게 한다고 밝혔습니다.
- Sundar Pichai는 NVIDIA와 HuggingFace가 파트너 관계를 이어가며 오픈 모델 생태계를 강화할 것이라고 썼고, Stability AI의 Emad Mostaque는 Hugging Face가 Stable Diffusion과 후속 모델의 출시·확장에 기여했다고 평가했습니다. 일부 포스트는 Claude·ChatGPT·Grok 장애와 소유권 문제를 연결했지만, 장애 원인이나 인수 조건은 원문에 구체적으로 나오지 않았습니다.
원문 트윗 2개 보기

Aravind Srinivas
A repository of open models and tools to train and serve them in an accessible manner is absolutely necessary for AI to remain accessible and useful to the public. Glad that NVIDIA is coming forth to do that for the open source community by supporting HuggingFace.
Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you
Congrats @JensenHuang @ClementDelangue for this great outcome! We were investors in Huggingface and continue to be partners and this will strengthen the open model ecosystem!
📈 GLM-5.3-Flash의 로봇 제어와 사용량 확대포스트 3
GLM-5.3-Flash가 별도 Robotics Fine-tuning 없이 카메라 프레임과 로봇 SDK를 이용해 이동·관절·그리퍼 동작을 조정한 사례가 공유됐습니다. 모델의 빠른 추론과 Coding Plan 사용량 확대가 함께 부각됐습니다.
세부 내용 보기
- Sentdex는 GLM-5.3-Flash가 카메라 프레임을 읽고 SDK의 이동·팔 관절·그리퍼 명령을 조합해 XGO mini 사족보행 로봇의 작업을 수행했다고 설명했습니다. 이 사례에는 해당 작업을 위한 학습이나 Fine-tuning이 없었고, 모델 제작 목적도 Robotics가 아니었다고 밝혔습니다.
- 기존 VLA 접근은 시뮬레이터 학습 파이프라인, Teleoperation 데이터, sim2real 문제와 작업·카메라·조명별 예외 처리 부담을 안지만, 해당 사례는 실제 환경과 작업별 학습 데이터 없이 장난감 집기 작업을 수행했습니다. GLM-5.3-Flash는 출시 일주일 뒤 기본 모델로 만들기 위한 개선점을 묻는 포스트와 ZCode 무제한 사용량·다른 에이전트 2배 quota 안내로 사용 확대도 함께 다뤄졌습니다.
원문 트윗 2개 보기
Harrison Kinsley
One question that's been on my mind for years now is: could we use regular multimodal LLMs not necessarily trained for robotics to do the high level robotics intelligence part that VLAs and WAMs attempt to do? The latest explosion of powerful opensource multi-modal LLMs has, IMO, begun to make this possible due both to intelligence and speed. This is GLM 5.3 Flash, which has vision understanding, but isn't meant to be a VLA/VLM/WAM/robotics model at all, controlling an XGO mini wheeled robot quadruped with an arm & gripper. GLM 5.3F simply has access to the robot's high level SDK for controlling movement, arm joints, open/close gripper...etc. It analyzes the frames from the camera and makes adjustments all on its own to solve the task. Nothing was trained here, nothing fine-tuned for this task. Z AI did not make this model for robots and tbh I think they're surprised this works when I talk to them about it! This also works quite well with DSV4F + a vision capable model like Qwen 3.8 27B. I havent tried JUST Qwen 3.8 27B, but I'm sure it works too. I like the "logic" to be a model that's as fast as possible (but still intelligent). There's also an experimental vision version of DSV4F, I'm confident that'll work too and might even be better bc the full loop might be the fastest of all with this model. An obvious question you might wonder is: well why not use VLA or VLM? The hard part about robotics isn't object detection, that's long solved. This also isn't a solution for gait/locomotion...yet, but I actually don't think this is far away either and I've done some experimentation with LLMs in this space in the past and it does show promise. It might actually already be here for quadrupeds, since you dont need super fast IMU readings to maintain balance. I've also tried many of the larger, more generalist, VLAs that you should be able to use with popular robots and tbh there are just so many edge cases that make things hard and not work. You gotta get the camera, lighting, task, everything *just right* or the demo fails. This is for the actual hard part in robotics right now: intelligence, logic, and planning for all the ways the real world just simply isn't perfect. I've trained VLAs. They're super finicky and you're always running into sim2real issues, especially around the camera. You also have to build the whole training pipeline in a simulator, and, if everything does work, you still just have a robot that does this 1 single thing after weeks of work. If you use teleop, this overcomes the "2real" problem, but now you need to painstakingly collect teleop data, and it's only good at that specific task and that particular robot. There is a growing set of egocentric training data for "general purpose" VLAs and world action models (for humanoid form factors), but I'm really starting to wonder: Why? I think we might just sidestep this whole area of research entirely. I didn't need any training data or special environment to work with this quadruped and arm to do the task I was after. This particular quadruped and arm doesn't even exist in the wild yet really, it's a demo build from a company launching it on kickstarter, so it's not like this robot's data exists in the LLM to any real extent. I think this is cool as heck that this works and I am interested to see just how far I can push it. Also this marks the first time that I've finally got a generalist solution to a task I've been trying to solve ever since I became a dad of twins: pick up toys off the ground. This is a big day!

Zixuan Li
GLM-5.3-Flash has been live for a week. If we only improve one thing next, what would make it your default?
📈 WeatherNext 3의 전 지구 강수 예측 고도화포스트 4
WeatherNext 3가 실시간 관측을 직접 학습해 더 세밀한 전 지구 기상 예측을 빠르게 수행하고, 강수 예측 오차를 최대 50% 줄였다는 수치가 공유됐습니다. Google Search와 GeminiApp 등 서비스 적용도 시작됐습니다.
세부 내용 보기
- 기존 전 지구 기상 모델이 강수 분포를 흐리게 만들거나 강한 폭풍 경계를 놓치는 문제가 있었지만, WeatherNext 3는 실제·실시간 관측을 직접 학습해 지역화된 예측을 더 빠르게 산출하는 구조로 소개됐습니다. GoogleDeepMind는 전 지구 강수 예측에서 오차를 최대 50% 줄였다고 밝혔습니다.
- Brightband 실시간 leaderboard에서 최상위 전 지구 기상 모델이라는 평가와 함께 고해상도 기온·강수 예측이 공유됐고, Google Search·GeminiApp·Google Maps·Weather API·Google Earth Engine에 적용되기 시작했습니다. 게시물에는 절대 오차값이나 비교 모델명은 추가로 나오지 않았습니다.
원문 트윗 2개 보기
Google DeepMind
WeatherNext 3 is a major breakthrough in how we forecast global weather. ⛅ Developed with @GoogleResearch , the model learns directly from real-world, real-time observations to give more localized highly accurate predictions faster. 🧵
Predicting rain accurately is notoriously difficult for global weather models, with previous methods producing blurry estimates or missing severe storm boundaries. WeatherNext 3 achieves a major leap in global precipitation forecasting, delivering up to a 50% reduction in error
➖ AI 생성 이미지 판별을 위한 Percept-Lens포스트 1
Percept-Lens는 범용 vision model을 고정한 뒤 특징 공간에 Gaussian decision rule을 적용해 실제 이미지와 AI 생성 이미지를 판별했습니다. 익숙하지 않은 생성기·프롬프트·스타일·도메인에서 탐지기가 약해지는 문제를 평가 대상으로 삼았습니다.
세부 내용 보기
- 기존 AI 생성 이미지 탐지기는 익숙한 이미지에서는 작동해도 생성기·프롬프트·스타일·이미지 도메인이 바뀌면 성능이 급격히 떨어질 수 있습니다. Percept-Lens는 이 변화 조건을 평가하는 공통 benchmark를 만들고, 탐지기 실패가 vision model의 구분 능력 상실인지 결정 규칙의 실패인지 나눠 확인했습니다.
- 새 탐지기는 범용 vision model의 가중치를 바꾸지 않고, 특징 공간에서 라벨이 있는 실제 이미지와 AI 생성 이미지의 분포를 Gaussian decision rule로 모델링한 뒤 새 이미지를 더 가까운 집단에 배정합니다. 게시물은 같은 광범위한 평가 세트에서 기존 공개 탐지기 가운데 가장 강한 모델보다 높은 성능을 기록했다고 밝혔지만, 구체적인 점수는 공개하지 않았습니다.
➖ 에이전트 실행 계층의 저장소·도구·운영 제어포스트 4
에이전트 제품의 관심이 모델 자체에서 저장소 교체, MCP 연결, 자체 머신 실행, quota·예산·관측성 관리로 이동했습니다. LangChain과 MongoDB의 BackendProtocol, Cloudflare의 자체 인프라 실행, AI Gateway의 운영 기능이 한 흐름으로 묶였습니다.
세부 내용 보기
- LangChain은 stateless MCP 규격을 기본 패키지로 옮기고 interrupt 기반 elicitation과 list endpoint caching을 추가했습니다. Deep Agents는 read·write·glob·grep 호출을 유지한 채 BackendProtocol 뒤의 저장 계층을 교체하며 MongoDB Atlas를 연결할 수 있게 됐습니다.
- Cloudflare와 Cursor 조합은 Cursor의 에이전트 loop를 유지하면서 수요에 따라 확장되는 자체 머신과 내부 서비스·특수 하드웨어를 연결하고, Vercel AI Gateway는 여러 coding agent를 한 계층으로 묶어 uptime·관측성·예산 관리와 전환 편의성을 제공하는 방식입니다. LangChain은 에이전트와 평가 모델을 같은 계열로 고르면 mode collapse가 생길 수 있어 다른 모델 계열을 써야 한다는 운영 규칙도 공유했습니다.
원문 트윗 2개 보기
Sydney Runkle
we now support the new stateless MCP spec in @LangChain ! this revamp includes moving MCP support into the main langchain package, supporting elicitation via interrupts, list endpoint caching, and more!

Harrison Chase
Agent workspaces should be durable, inspectable, and swappable Deep Agents’ virtual file system is intentionally behind a BackendProtocol, so the agent code can keep using read/write/glob/grep while teams choose the storage layer that fits production Nice to see MongoDB support this!
MongoDB now powers the virtual file system behind @LangChain's Deep Agents. 🚀 This new integration lets you implement Deep Agents' BackendProtocol against MongoDB Atlas instead of building your own storage layer. And agent code calling read_file, write, glob, and grep don’t
➖ 실시간 Video Generation과 World Language Action Models포스트 1
Reka와 NVIDIA가 30B 파라미터 Video Generation 모델을 실시간 버전으로 구축했습니다. 자연어로 생성 중간에 조정하고 단일 H100에서 720p·24fps를 11.8배 빠르게 처리하는 구성이 핵심입니다.
세부 내용 보기
- 기존 Video Generation이 결과물을 만든 뒤 끝나는 흐름과 달리, 이 모델은 생성 도중 자연어 지시로 방향을 바꿀 수 있도록 설계됐습니다. Reka는 30B 파라미터 모델을 720p·24fps로 실행하고 단일 H100에서 11.8배 속도 향상을 기록했다고 밝혔습니다.
- blind evaluation에서 큰 품질 저하가 없었다는 결과와 함께 closed beta가 열렸고, 물리 세계를 생성하는 데서 그치지 않고 그 안에서 행동하는 World Language Action Models로의 확장이 목표로 제시됐습니다. 게시물에는 평가 세부 데이터나 비교 기준은 나오지 않았습니다.
용어 해설
- 자율 에이전트(Autonomous Agents)
- — 외부 지시를 계속 기다리지 않고 목표에 따라 판단과 행동을 반복하는 AI 시스템입니다. 자율성이 높아질수록 중요 프로세스의 통제권과 처리 과정을 사람이 확인할 수 있는 구조가 필요합니다.
- 오픈 모델(Open Models)
- — 가중치나 개발 생태계가 공개되어 개발자·연구기관·기업이 직접 활용하거나 목적에 맞게 수정할 수 있는 AI 모델입니다. 접근성과 사용자 주도권을 높이는 기반으로 거론됩니다.
- 고정 표현(Frozen Representations)
- — 이미 학습된 모델의 내부 특징 표현을 추가 학습으로 바꾸지 않고 그대로 사용하는 방식입니다. Percept-Lens는 이 표현 공간에서 실제 이미지와 AI 생성 이미지의 분포를 비교해 분류합니다.
- MCP
- — AI 애플리케이션과 외부 도구·데이터 소통을 위한 표준 규격입니다. LangChain의 새 구현은 상태를 저장하지 않는 규격과 중단을 통한 사용자 입력, 목록 endpoint 캐시를 지원합니다.
- 가상 파일 시스템(Virtual File System)
- — 에이전트가 read·write·glob·grep 같은 파일 연산을 일관된 인터페이스로 수행하는 저장 계층입니다. BackendProtocol을 두면 실제 저장소를 MongoDB Atlas 등으로 교체할 수 있습니다.
- 강수량 예측(Precipitation Forecasting)
- — 비와 같은 강수의 위치와 양을 예측하는 기상 모델 작업입니다. WeatherNext 3는 실시간 관측을 직접 학습해 전 지구 강수 예측 오차를 최대 50% 줄였다고 발표됐습니다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.