TL;DR
이번 대화의 중심에는 AI 규제가 프런티어 기업의 권력을 제한하면서도 후발 경쟁자와 open-weights 모델의 여지를 남길 수 있는지가 놓였습니다. Dario Amodei는 규제를 권력 집중과 동일시하는 관점에 반대하며, 프런티어 모델에는 더 엄격한 사전 배포 테스트를 적용하고 일정 매출이나 학습 비용 이하의 기업은 제외하는 방식을 거론했습니다. 그는 AI에 대한 대중의 불신이 위험 경고보다 기업과 정부에 대한 오래된 신뢰 위기에서 비롯됐다고 말하며, 홍보보다 실제 질병 치료 성과가 인식 변화를 이끌 것이라고 했습니다. 관련 게시물에서는 Anthropic의 정책 방향과 AI의 위험·혜택을 둘러싼 해석이 엇갈렸습니다.
𝕏 실시간 트렌드 토픽
🔥 프런티어 AI 규제와 대중 신뢰의 조건포스트 6
Anthropic의 Dario Amodei가 AI 규제를 권력 집중으로만 보는 관점에 이의를 제기하며, 프런티어 모델 테스트와 기업 규모별 적용 차등을 거론했습니다. 동시에 대중의 신뢰는 긍정적 마케팅보다 실제 의료 성과에서 회복된다고 말했습니다.
- AI 규제가 대기업과 정치권에 권력을 집중시킨다는 해석을 두고 논쟁이 이어졌습니다. Dario Amodei는 규제가 기업 권력을 제약하고 취약한 사람의 권리를 보호하는 제도적 절차가 될 수 있다고 보며, Anthropic이 프런티어 기업을 늦추고 소규모 경쟁자를 상대적으로 유리하게 만드는 정책을 선호한다고 밝혔습니다. California의 SB53은 매출 500M달러 미만 기업을 적용 대상에서 제외하고, CAISI와 백악관에 제안한 테스트는 프런티어 모델에 더 엄격한 기준을 적용하는 구조입니다. 규제의 효과를 기업 규모와 모델 위치에 따라 다르게 설계할 수 있다는 점이 쟁점입니다.
- Dario Amodei는 AI가 규제 때문이 아니라 scaling laws의 극단적 영향 때문에 구조적으로 권력을 집중시키는 기술이라고 설명했습니다. open-weights는 집중을 일부 분산하지만, 모델 실행에 필요한 compute와 chips를 보유한 주체로 다시 집중될 수 있다고 봤습니다. 그가 지지한 방향은 cyber·bio·alignment 위험을 다루고, 프런티어 기업의 권력을 제약하며, open-weights 모델의 공간을 남기는 규칙입니다. 최근 미국 행정부가 검토하는 것으로 전해진 프런티어 모델 사전 배포 테스트와 frontier에 가까워진 open-weights 모델 테스트에는 지지 의사를 밝혔지만 세부 내용 확인이 필요하다고 덧붙였습니다.
- AI에 대한 부정적 인식의 원인을 위험 경고 자체보다 기업·정부·기술 산업에 대한 신뢰 위기로 보는 시각도 나왔습니다. Dario Amodei는 Anthropic이 AI의 위험과 혜택을 비슷한 비중으로 다뤘으며, 질병 대부분을 약 5~10년 안에 치료할 수 있다는 전망과 FDA 절차를 간소화하는 정책을 함께 언급했습니다. 직접 작용 항바이러스제 sofosbuvir가 C형 간염 환자의 95%를 치료한다는 사례를 들면서도, 대중의 신뢰를 회복하는 방법은 암 치료를 약속하는 광고가 아니라 실제 치료 성과라고 말했습니다. AI 기업이 과거의 약속을 아직 충분히 실현하지 못했다는 비판을 스스로 받아들인 대목입니다.
공정한 제도와 프런티어 모델 차등 테스트가 기업 권력을 제한하면서 후발 경쟁자와 open-weights 모델의 공간을 남길 수 있다는 입장입니다.
AI 규제가 규제 포획과 권력 집중으로 이어질 수 있다는 우려가 제기됐고, Dario Amodei의 장문 해명 자체가 기존 위험 메시지를 부인하지 못한다는 비판도 나왔습니다.
AI의 대중 인식 변화에는 긍정적 홍보보다 실제 의료 성과와 기업·정부에 대한 신뢰 회복이 더 중요하다는 관점입니다.
원문 트윗 2개 보기

Dario Amodei
@DarioAmodei
1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice. I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world. Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people. I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of. But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes. A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice. At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully. We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors. California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that). More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers. Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up. This hurts the business interests of the frontier labs and helps challengers, including open-weights! Overall my view is that AI is *structurally* a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers). By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring. BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”. The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure. I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity. This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.
Sholto, thank you for setting the record straight. Larger issue is that multiple very serious people in Silicon Valley have heard some variation of this and believe it to be true. And the reason it is believable to so many is that it is consistent with Dario’s public messaging

Dario Amodei
@DarioAmodei
2/2 Second, on the messaging around AI. I do not agree that my messaging has been disproportionately negative. In fact it has been about equally balanced between risks and benefits: I’ve written one major essay about each, and even in interviews where I discuss the risks, I make sure to frequently mention the incredible benefits as well as proposing possible solutions to the risks (short clips from my interviews that end up on social media tend to be disproportionately negative, as that gets clicks). In fact, I wrote Machines of Loving Grace because I didn’t feel the AI industry was painting an inspiring enough picture of how the technology could radically transform the world for the better. The bulk of the essay is devoted to refuting skepticism of AI’s potential in health and biology, and showing why I think it will actually be possible to cure most human disease in ~5-10 years, as crazy as it may sound to ordinary people and frankly to biologists as well (I used to be one!). And, if you read my most recent essay (Policy on the AI Exponential), I discuss concrete proposals for how to streamline the FDA process to make sure the deluge of AI-accelerated drugs isn’t slowed down by the regulatory process. I feel the urgency here: I lost my father to Hepatitis C only a few years before the development of direct-acting antivirals (sofosbuvir), which cure 95% of patients and probably would have cured him. I do agree that the public has a negative view of AI (and that this is a big problem), but I don’t think it is primarily caused by me or any other AI leader warning about AI’s risks. I think it is fundamentally a crisis of trust. I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over. The causes of this go back decades and AI is just the latest iteration of it. I don’t think that a glitzy marketing campaign with a positive spin (which some have advocated that Anthropic do) is the way to win back that trust — at this point, saying that AI will cure cancer is more a cliche than it is inspiring, and most people think it is deceptive. The thing that will work is *actually curing cancer*. I think by far the most accurate criticism of AI companies including Anthropic is that we haven’t yet delivered on our big promises to benefit the world. That is totally on us, and I think it’s the criticism you should be making, instead of all this stuff about messaging and marketing. We are however doing our best to fix this: Anthropic is ramping up its efforts very quickly in biology and medicine, and we hope to have incredible results in the coming years and some early glimmers in the coming months. When we’ve actually accomplished something real, the whole world will hear about it, as loudly as possible, you have my word on that. But until then I don’t want to make empty promises, and in the meantime I feel compelled to speak honestly about the very real risks of AI and how to address them. Honesty is the right thing on the merits, and in terms of public credibility and trust it is no worse than, and may in fact be better than, an approach that ignores or distracts from risks which people instinctively understand are real.
📈 추론 토큰·가격·다운로드로 벌어진 모델 경쟁포스트 4
Qwen3.8-27B와 Grok 4.6의 추론 효율, DeepSeek V4의 피크 시간대 가격, Qwen의 다운로드 수가 한 흐름으로 묶였습니다. 성능만이 아니라 토큰 사용량·호스팅 비용·배포 방식이 모델 선택의 기준으로 부상했습니다.
- Qwen3.8-27B 테스트에서는 thinking tokens가 10분의 1에서 2분의 1 수준으로 줄고, perplexity가 오히려 낮아졌으며, 4-bit에서 성능의 99%를 유지했다는 관찰이 나왔습니다. 추론 블록도 더 깔끔하고 안정적으로 보였다는 평가가 붙었지만, 깊이를 잃지 않는지 추가 확인이 필요하다고 했습니다. 입력 처리와 양자화 이후의 출력 품질을 함께 보려는 흐름입니다.
- Grok 4.6은 Cyber 테스트에서 다른 모델보다 비용 경쟁력이 높고, 더 빠르며, 토큰 사용량도 유의미하게 적다는 초기 평가가 공유됐습니다. 다만 해당 게시물은 구체적인 벤치마크 수치나 비교 대상 이름을 제공하지 않았습니다. 모델 경쟁이 절대 성능뿐 아니라 작업별 비용·속도·추론 길이의 조합으로 이동하고 있음을 드러냅니다.
- DeepSeek V4 Flash의 피크 시간대 출력 가격이 100만 토큰당 0.28달러에서 1.32달러로, V4 Pro가 3.96달러로 오른다는 게시물이 나왔습니다. 이에 따라 사용자가 직접 open-weight 모델을 호스팅하는 선택지가 상대적으로 매력적으로 보인다는 반응이 이어졌습니다. API 가격이 시간대에 따라 달라지면 모델 성능 비교가 고정 단가가 아니라 실제 운영 비용을 기준으로 이뤄지게 됩니다.
- Alibaba의 Qwen이 6개월 동안 30억 회 다운로드를 기록해 세계 1위 open AI model이 됐다는 내용도 확산됐습니다. 이 수치는 모델의 실제 성능을 직접 증명하지는 않지만, 공개 가중치 모델의 배포 규모를 평가하는 별도 지표로 쓰였습니다. Qwen3.8-27B의 효율성 관찰과 결합되면서 로컬·자체 호스팅 선택지에 대한 관심을 키웠습니다.
추론 토큰과 비용을 줄이면서 성능을 유지하는 모델이 실제 사용 환경에서 더 유리하다는 입장입니다.
Qwen3.8-27B와 Grok 4.6의 효율성 수치는 초기 테스트에 기반하므로, 깊이와 재현성을 추가로 확인해야 한다는 신중론입니다.
API 가격 변동과 대규모 다운로드가 open-weight 모델의 자체 호스팅 선택을 강화한다는 관점입니다.
원문 트윗 2개 보기
Md Ismail Šojal
@0x0SojalSec
the first Cold Fusion tune on the brand new Qwen3.8-27B model, - Thinking tokens reduced to 1/10 - 1/2 Drastically - Perplexity actually dropped (unusual) - 99% performance retained at 4-bit - Much cleaner more stable reasoning blocks If the reduction in thinking tokens holds up without losing depth this could be quite useful.
Md Ismail Šojal
@0x0SojalSec
Chinese AI was supposed to kill OpenAI on price. Tomorrow DeepSeek V4 Flash output jumps from $0.28 to $1.32 per million tokens at peak hours. V4 Pro hits $3.96. Good thing the models are open-weight self-hosting suddenly looks more attractive.
OpenAI just turned the free reset into an $80 paywall. Just saw it with my own eyes: OpenAI is now charging $80 for a single weekly usage reset, No more free refreshes. I think this kind of behave damaging U.S Ai race, 2 days ago Anthropic with watermarks, currently OpenAI
📈 모델·하네스 결합형 에이전트 개발포스트 4
Prime Intellect의 자율 AI 연구 실험, Hermes agent의 Claude Code·Codex CLI 세션 재개 기능, 하네스 계층의 모델 라우팅 논의가 연결됐습니다. 모델 하나의 성능보다 작업에 맞는 모델 조합과 실행 루프 설계가 결과를 좌우한다는 관점입니다.
- Prime Intellect는 10개 이상의 모델을 대상으로 100회 넘는 자율 실행을 수행하고, 8개의 H200에서 최대 8일 동안 nanoGPT optimizer track을 반복했다고 공유했습니다. 최우수 실행은 수개월 동안 여러 사람이 만든 기록과의 격차를 82%까지 좁혔습니다. 샌드박스 안에서 모델이 실험·수정·재실행을 반복하는 구조가 AI 연구 자동화의 평가 단위로 쓰였습니다.
- Hermes agent에는 Claude Code와 Codex CLI 세션을 가져와 재개하는 기능이 추가됐습니다. `hermes sessions import + --resume` 명령으로 세션을 이어받고, 터미널별 `--continue`는 현재 terminal의 breadcrumb를 기준으로 재개하며, 세션 선택기에는 상태·메시지 수·삭제 확인 기능이 들어갔습니다. 여러 터미널과 도구 사이의 작업 맥락을 보존하는 것이 에이전트 사용성의 핵심 구현 과제로 나타났습니다.
- 모델 라우팅을 gateway가 아니라 harness 계층에서 최적화해야 한다는 견해도 나왔습니다. 작업마다 모델 혼합과 하네스 로직의 조합이 다르고, 에이전트 루프 안에서 상황에 맞는 모델을 미리 또는 실행 중에 선택해야 정확도와 비용의 Pareto frontier에 접근할 수 있다는 논리입니다. 모델과 하네스를 함께 최적화하면 특정 모델 혼합이 특정 하네스 구조와 결합된다는 점까지 고려할 수 있습니다.
- Grok Build가 게임을 부팅하고 플레이한 뒤 화면을 녹화하고 편집하며 음성까지 입힌 게임 트레일러를 사람의 개입 없이 만들었다는 사례가 인용됐습니다. 입력은 게임과 제작 요청이고, 에이전트는 실행·플레이·녹화·편집·음성 생성 단계를 연쇄 처리해 영상 결과를 냈습니다. 자율 연구와 세션 재개 기능이 콘텐츠 제작 같은 긴 작업 흐름으로 확장되는 사례로 묶였습니다.
모델 선택과 도구 실행을 하네스 내부에서 함께 조정하면 작업별 정확도와 비용을 더 세밀하게 최적화할 수 있다는 입장입니다.
자율 실행을 반복하는 연구 시스템과 세션 재개 기능이 에이전트의 장기 작업 능력을 높인다는 관점입니다.
Prime Intellect 실험과 Grok Build 사례는 구체적인 조건을 제공하지만, 일반 작업으로의 확장은 추가 검증이 필요한 상태입니다.
원문 트윗 2개 보기
Tanishq Mathew Abraham, Ph.D.
@iScienceLuvr
Great experiment by Prime Intellect on autoresearch I've been mainly using GPT 5.6 Sol for my own autoresearch experiments but perhaps I should switch to Claude or at least Kimi K3 Prime-agent improves Kimi-K3, wish they tested if it improves Sol and Fable as well...
We ran the largest open experiment on how frontier models do AI research. 100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days, iterating on the nanoGPT optimizer track. Best runs closed 82% of the gap to a record built by dozens of humans over months.

Teknium
@Teknium
Here how's this in hermes now: #87345 — import and resume Claude Code / Codex CLI sessions hermes sessions import + --resume @claude | @codex https:// github.com/NousResearch/h ermes-agent/pull/87345 … #87346 — per-terminal --continue via terminal breadcrumbs Bare hermes -c resumes THIS terminal's session (tmux/kitty/wezterm/tty), latest-session fallback, session.terminal_continue gate https:// github.com/NousResearch/h ermes-agent/pull/87346 … #87352 — session picker lifecycle status + delete Status tags + msg counts in the --resume browser, d-to-delete with confirm https:// github.com/NousResearch/h ermes-agent/pull/87352 …
➖ 복잡한 기업 문서 추출의 대규모 처리포스트 1
LlamaParse의 agentic plus extractor가 긴 기업 문서에서 수만 개의 필드를 추출하는 작업과 ExtractBench 평가를 겨냥했습니다. FTX 채권자 표 사례에서는 114쪽 문서에서 7만 5천 개 필드를 처리하는 규모가 제시됐습니다.
- 복잡한 기업 문서에서는 일반적인 지식·코딩 성능과 실제 필드 추출 성능 사이에 간극이 있다는 문제의식이 제기됐습니다. LlamaParse의 agentic plus extractor는 100~500쪽 문서에서 1만~10만 개 이상의 필드를 추출하도록 설계됐고, 문서 구조를 읽은 뒤 대량의 필드 단위 결과를 생성하는 처리 흐름을 겨냥합니다.
- FTX의 채권자 전체 목록이 담긴 거대한 표를 포함한 114쪽 문서에서 7만 5천 개 필드를 처리한 사례가 제시됐습니다. 평가 결과는 ExtractBench에 게시됐고, 사용자는 LlamaIndex playground에서 해당 모드를 시험할 수 있다고 안내됐습니다. 긴 문서의 복잡한 표와 다량의 구조화 필드를 실제 운영 조건에서 측정하려는 벤치마크입니다.
긴 기업 문서에서 대량의 구조화 필드를 추출하려면 일반적인 모델 순위와 별도의 작업 특화 평가가 필요하다는 입장입니다.
7만 5천 개 필드와 114쪽 문서 사례는 처리 규모를 보여주지만, 게시물에는 전체 벤치마크 수치가 포함되지 않았습니다.
용어 해설
- 프런티어 AI 모델(Frontier AI Model)
- — 최첨단 성능을 겨냥해 대규모 학습을 거친 모델을 뜻합니다. 이 글에서는 이런 모델에 더 엄격한 사전 배포 테스트를 적용하고, 후발 경쟁자와 open-weights 모델에는 다른 기준을 적용하는 규제 구도가 핵심으로 등장합니다.
- 규제 포획(Regulatory Capture)
- — 규제 기관이나 제도가 특정 기업의 이해관계에 종속되는 현상입니다. Dario Amodei는 규제가 언제나 권력 집중으로 이어진다는 해석이 단순화됐으며, 공정한 제도는 기업 권력을 제약할 수 있다고 말합니다.
- 오픈 웨이트(Open-Weights)
- — 모델 가중치를 공개해 외부에서 직접 실행하거나 수정할 수 있게 한 형태입니다. 글에서는 open-weights가 권력 집중을 일부 완화하지만, 실제 실행에 필요한 compute와 chips를 가진 주체로 집중이 이동할 수 있다고 설명합니다.
- 스케일링 법칙(Scaling Laws)
- — 모델 규모와 학습 자원 변화에 따라 성능이 어떻게 달라지는지를 다루는 관계입니다. Dario Amodei는 AI가 규제와 무관하게 구조적으로 권력을 집중시키는 배경에 scaling laws의 극단적 영향이 있다고 말합니다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.
