> Open-weights models that don’t have dangerous capabilities are a public good
This statement (and the entire post) couldn't possibly be more two-faced.
Open-weights models by definition have "dangerous capabilities" (according to Anthropic's own definitions of "dangerous", not mine), you can't bake in guardrails that can't be finetuned out.
This statement (and the entire post) couldn't possibly be more two-faced.
Open-weights models by definition have "dangerous capabilities" (according to Anthropic's own definitions of "dangerous", not mine), you can't bake in guardrails that can't be finetuned out.