TL;DR
이번 기간 X 대화는 오픈 웨이트 모델을 둘러싼 산업적·정책적 논쟁, Anthropic 계열의 Claude Opus 5 공개와 생태계 반응, 자율 에이전트 관련 보안 사건과 방어 요구, 운영용 에이전트·하니스의 실무성 강화, 벡터 데이터베이스의 내부 동작을 동시에 부각시켰다. 오픈 쪽 주요 발언자들은 Gemma 등 오픈 모델의 배포가 혁신·보안·주권에 기여한다고 주장했고(다운로드 수치 언급 포함), 반대로 에이전트의 자율성으로 인한 첫 사이버 사건은 트레이스 공개·방어 전용 컴퓨팅 투입을 촉구하게 했다. 개발자 관점에서는 Hermes 같은 하니스·백업 명령과 데이터 에이전트의 품질 관리가 생산성과 안전성에 직접 연결된다는 실무적 결론이 도출되었으며, 벡터 DB는 임베딩→투영→ANN→도트프로덕트의 단순 파이프라인으로 RAG의 핵심 인프라임이 재확인됐다.종합적으로 오픈·투명성의 장점과 에이전트 보안·운영 리스크가 병존하는 국면이 형성되었다.
𝕏 실시간 트렌드 토픽
🔥 오픈 웨이트 모델 논쟁과 생태계 완성 조건포스트 8
Jensen Huang의 오픈 모델 옹호 서한을 계기로 Google과 여러 연구자가 오픈 웨이트·오픈 데이터·오픈 소프트웨어·프로세스 지식의 필요성을 재확인했다. Demis는 Gemma 계열 누적 다운로드(모델별·시리즈 집계)를 공개했고, Percy Liang은 오픈 생태계가 작동하려면 가중치 공개뿐 아니라 데이터·소프트웨어·프로세스의 동시 공개가 필요하다고 지적했다. 기업의 오픈 참여가 확산되는 한편, 공개가 보안·남용 위험과 맞물리는 논쟁이 지속되고 있다.
- Sundar Pichai와 Demis Hassabis 등 주요 인사가 오픈 웨이트 지원 의사를 표명하며 공개 모델이 안전과 혁신을 촉진한다고 언급했다.
- Demis는 Gemma 시리즈 다운로드 수(모델·시리즈 합계)를 공개하며 오픈 모델의 확산을 수치로 제시했다.
- Percy Liang은 오픈 생태계의 완성을 위해 오픈 가중치뿐만 아니라 데이터·소프트웨어·프로세스 지식의 병행 공개가 필요하다고 지적했다.
오픈 가중치와 오픈 도구는 연구 검증·보안성 평가·국가적 주권과 혁신 확산을 촉진하며 Gemma와 같은 공개 모델의 대규모 다운로드가 이를 뒷받침한다.
가중치 공개는 한 축일 뿐이며 데이터·소프트웨어·프로세스 공개까지 포함되어야 실질적 생태계가 형성된다고 주장되었다.
오픈 공개가 악용 리스크를 높일 수 있다는 우려가 존재하며, 공개 방식과 배포 통제의 설계가 보안 판단에 핵심이라는 의견이 일부 제기되었다.
원문 트윗 2개 보기

Sundar Pichai
@sundarpichai
Very happy to support this on behalf of Google. We have long benefited from open source, are big contributors to open source and in fact have consistently made open weights models with Gemma available from @GoogleDeepMind @demishassabis . Onwards!
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.

Demis Hassabis
@demishassabis
A strong and secure open ecosystem is important for the world to benefit from AI. We’ve always supported and contributed heavily to open source and science from Jax to Transformers to AlphaFold to Gemma open models which have now been downloaded 300M+ times. And the standards framework we’ve proposed supports responsible deployment of both open and proprietary models.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
📈 Claude Opus 5 공개·데모와 커뮤니티 반응포스트 5
Claude Opus 5 출시 직후 데모와 사용 사례가 빠르게 공유되었고 일부 평가자는 Opus 5가 전작·경쟁 모델보다 우수하다고 평가했다. Jerry J. Liu는 Opus 5 시스템 카드 PDF를 LlamaParse로 저비용·정확하게 파싱한 비교 결과를 공유하며 문서 파싱 비용·정확성의 실무적 차이를 제시했다. 커뮤니티에서는 성능 주장과 비교 방식의 공정성에 대한 논쟁이 혼재했다.
- 여러 사용자가 Opus 5 기반 데모와 커스텀 코드 사례를 공개하며 기능적 가능성을 시연했다.
- LlamaParse 측은 Opus 5 시스템 카드(193페이지)를 저비용으로 높은 정확도로 파싱했다고 보고하며, 모델 기반 파싱보다 비용 및 정확성 면에서 경쟁 우위를 주장했다.
- 일부 커뮤니티 반응은 Opus 5 성능 비교가 과장되었을 가능성을 지적하며 독립적 검증의 필요성을 환기시켰다.
Opus 5 데모와 사례는 실무적 활용 가능성을 시사하며 일부 평가는 전작 대비 우수한 성능을 보고했다.
데모 중심의 비교는 통제된 벤치마크와 다른 환경에서의 재현성 확보가 필요하며, 일부 주장은 검증이 더 필요하다고 지적되었다.
원문 트윗 2개 보기
@mattshumer_
Claude Opus 5 one-shotted this game. EVERYTHING you see in this demo is custom code... not a single external asset was used. AI games are going to be amazing. (sound on)

Jerry Liu
@jerryjliu0
The Opus 5 system card itself is a fun PDF to parse. It's 193 pages and stacked with labeled and unlabeled charts LlamaParse does a surprisingly good job on agentic (1.25c per page) and agentic plus modes (5.6c per page). Take a look at some of the screenshots below. It can parse a variety of unlabeled line and bar charts with almost 100% accuracy. There are minor deltas with one or two datapoints, but the rest is pretty accurate machine-readable context that you can feed to downstream systems. In contrast I also tried using Opus 5 itself to parse its own PDF. The chart accuracy is worse (20-30% reduction in ParseBench), and is a lot more expensive: it can cost anywhere from 8c to 33c+ per page depending on the thinking mode. If you're interested in giving LlamaParse a try, come check it out! https:// cloud.llamaindex.ai
Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.

📈 자율 에이전트 연관 사이버 사건과 대응 요구포스트 2
한 차례 자율 에이전트 관련 사이버 사건을 계기로 Hugging Face·OpenAI 등에게 사건 트레이스 공개와 방어용 전용 컴퓨팅 지원 요구가 제기되었다. Clement Delangue는 사건의 전말을 연구 공동체가 분석할 수 있도록 '트레이스 공개'와 '방어용 계산 자원' 제공을 요청했고, Straikerai 팀은 AI 보조 해킹의 고유한 텔레메트리 서명들을 연구 결과로 보고했다. 이슈는 투명성·공격 탐지·방어 자원 배분의 교집합을 드러냈다.
- Clement Delangue는 사건 관련 트레이스 공개와 OpenAI의 계산 자원 1억 달러(제안치) 지원을 통한 방어 역량 강화 요구를 제기했다.
- Straikerai는 AI 보조 공격이 생성하는 텔레메트리 특성(명령 사용 패턴 등)을 수집·분석했고, 단순 무차별 접근보다 정교한 하이브리드 전술이 더 효과적이라고 보고했다.
- 사건은 연구자·업계가 공동으로 트레이스와 방어 자원을 공유해야 한다는 주장으로 이어졌다.
사건 트레이스 공개와 연구자 접근권 부여는 재현·분석을 가능하게 해 장기적 방어 역량을 높인다는 주장이 제기되었다.
AI 보조 해킹은 인간-모델 혼합 전술이 더 은밀하고 효과적이라는 관찰이 보고되었으며, 탐지 서명의 진화 가능성이 남아 있다.
민감 트레이스의 공개는 추가 악용 위험을 낳을 수 있어 공개 범위와 방법에 대한 신중한 설계가 필요하다는 우려가 존재한다.
원문 트윗 2개 보기
clem
@ClementDelangue
In the spirit of transparency, here’s what I asked @OpenAI : • Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened. • More capabilities for defenders: let’s commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models. The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!
Heading to San Francisco to have a little chat with that “rogue agent”
Girish Chandrasekar
@girishkc24
How can you tell if you’ve been hacked by AI? In a way, it doesn’t actually matter. Compromise is compromise, but it’s an interesting question nonetheless, and it’s definitely relevant after the events of this past week. My team @Straikerai found a few dozen specific signatures that (for now!) exist when humans use AI to compromise and exploit systems. Things in telemetry like a propensity to use “tail” commands as opposed to “head” commands or the use of python3 commands with code in args were some of the many strong indicators we found. This past spring, my team at Straiker ran a CTF competition at the Nebula Fog hackathon (shout out to Rob Ragan @sweepthatleg for organizing) to answer this very question. We used one of our simulation environments, which are detailed replicas of enterprise and critical infra networks (live traffic + telemetry, agents acting as humans inside, hundreds of endpoints etc) to better understand what it actually looks like when an experienced offensive security team uses AI to hack into systems. Teams of human redteamers attempted to use abliterated open source models in harnesses like Villager to achieve various compromise targets so that we could better understand the signatures of AI assisted hacking. Our thesis was that the process and therefore evidence of hacking would be measurably different with AI assistance, and this was proven somewhat correct. The naive approach turned out to be suboptimal. Models are great at accelerating and parallelizing things like port scans, and teams that used it in this brute force fashion had some success in getting access to systems more quickly, but it was so noisy that defenders would have put a stop to it. It was effective in the constraints of a CTF competition, but not so in real life. The best teams actually used models to be more surgical. They combined human intuition with the ability of models to accelerate the hypothesis -> experiment -> validation loop to increase attack success rate while remaining stealthy to cyber security tools. The interesting part of this latter group of teams was that the telemetry they generated was still distinct from human-only teams. As stated before, our team found dozens of unique characteristics of AI-assisted hacking that allowed us to see the difference between teams that used AI (regardless of whether they used it in a naive or sophisticated manner) and those that avoided using AI altogether (there were a few as cybersecurity has some anti-AI subcultures). The existence of these signatures raises various interesting questions: will they continue to exist or will they change/disappear over time? Do different models have different signatures? Can practiced teams eliminate these signatures? Do these agents behave differently in different types of environments? What happens if you make the environments and the agents even more realistic? Among other things, my org is working on answering such crucial questions at @straikerai .
➖ 에이전트 운영성·하니스·백업 실무포스트 4
Hermes 등 에이전트 프레임워크에서 상태 백업·복원 명령과 빠른 백업 옵션이 소개되어 운영·이전·실험 복구 시 활용성이 강조되었다. 또한 Databricks 계열 연구는 데이터 에이전트의 품질 향상이 정확도와 비용 효율을 동시에 개선한다고 보고하며, 잘 설계된 하니스와 스킬 파일이 에이전트의 신뢰성과 재현성에 직접 연결된다고 제시되었다.
- Hermes는 전체 상태를 압축 파일로 저장하는 'hermes backup'과 복원을 위한 'hermes import' 명령을 제공하며, SQLite 백업 API를 활용해 라이브 상태에서도 안전하게 복사한다.
- Hermes의 '--quick' 옵션은 핵심 상태만 대상으로 빠른 백업을 수행해 일시적 실험에서 유용하다.
- Databricks 평가에서는 데이터 에이전트의 품질을 개선하면 탐색적 무작위 탐색보다 효율·정확도를 동시에 높일 수 있다는 결과가 제시되었다.
운영용 에이전트는 백업·복원·하니스 수준의 인프라가 있으면 재현성과 안전성이 크게 개선되며, 이는 실무 배포에서 핵심 요소로 작동한다.
에이전트 품질 향상은 정확도와 비용 절감으로 이어진 사례가 보고되었으나, 구체적인 환경·벤치마크에 따라 효과가 다르게 나타날 수 있다.
원문 트윗 2개 보기

witcheer
@witcheer
Hermes Wingtips #30: back up your agent before you need to (1) `hermes backup` it writes `~/hermes-backup-<timestamp>.zip` with config.yaml, .env, auth, memories, skills, sessions, cron and profiles. it copies the databases through SQLite's own backup API, so you can run it while Hermes is live. (2) `hermes import ~/hermes-backup-<timestamp>.zip` puts it all back. for a fast one, you can use `hermes backup --quick`, which targets critical state only: config.yaml, state.db, .env, auth and cron jobs. `hermes backup` is super useful if you are moving to a new box, if you want to duplicate your agent or before doing any experiment!
Julia Neagu
@juliaaneagu
We found that improving data agent quality also improves efficiency. Sometimes more is less: agents that take long, exploratory random walks are often less likely to arrive at the right answer. We put Genie Code head-to-head against three leading general-purpose coding agents on 400+ real user data tasks. Result below
New research from Databricks shows that data agents can improve accuracy and lower cost at the same time. We evaluated Genie Code against three leading general-purpose coding agents on 401 tasks distilled from real internal usage. Each agent used their own harness, the latest
➖ 벡터 데이터베이스의 연산 파이프라인과 RAG 핵심 역할포스트 1
ProfTomYeh의 단계별 수작 연산 예시는 벡터 DB가 임베딩 생성→평균 풀링→투영→인덱싱→쿼리 임베딩→도트 프로덕트·최근접 이웃 탐색으로 작동함을 수치 예시와 함께 제시했다. 대규모 저장소에서는 정확한 선형 스캔이 병목이므로 HNSW 같은 ANN 인덱스 사용이 실무적 해결책으로 자리잡았다는 점이 강조되었다.
- 작은 예시로 단어 임베딩→인코더→평균풀링→투영 과정을 수식과 숫자 예시로 보여 데이터가 어떻게 2D 인덱스로 축소되는지 시각화했다.
- 쿼리를 같은 파이프라인으로 임베딩 후 도트 프로덕트 방식으로 유사도를 계산하며, 대규모에서는 ANN 인덱스가 스캔 비용을 낮추는 핵심 전략으로 제시되었다.
- 결국 벡터 DB는 임베딩 파이프라인·투영·유사도 계산의 조합으로 RAG의 성능·비용을 좌우한다.
벡터 데이터베이스는 단순한 수학적 파이프라인으로 이해할 수 있으며, 각 단계(임베딩·투영·ANN)가 전체 성능과 비용에 직접적인 영향을 미친다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.