AI Coding Community

368 bookmarks
Custom sorting
(13) OpenAI on X: "We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work. As the threat landscape evolves, we’re putting frontier intelligence in the hands of trusted defenders before attackers can deploy https://t.co/6o3GtxCxRA" / X
(13) OpenAI on X: "We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work. As the threat landscape evolves, we’re putting frontier intelligence in the hands of trusted defenders before attackers can deploy https://t.co/6o3GtxCxRA" / X
·x.com·
(13) OpenAI on X: "We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work. As the threat landscape evolves, we’re putting frontier intelligence in the hands of trusted defenders before attackers can deploy https://t.co/6o3GtxCxRA" / X
Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.
Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.
·x.com·
Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.
(20) Povilas Korop | Laravel & AI Coding Educator on X: "Tencent Hy3 model is surprisingly good for coding, for its price. Tested my new LLM benchmark projects on it, with OpenCode Go. Pleasantly surprised. P.S. v2 leaderboard is "work in progress", will finish it this week and will shoot a video probably on weekend or next week. https://t.co/yZSE6NUSgU" / X
(20) Povilas Korop | Laravel & AI Coding Educator on X: "Tencent Hy3 model is surprisingly good for coding, for its price. Tested my new LLM benchmark projects on it, with OpenCode Go. Pleasantly surprised. P.S. v2 leaderboard is "work in progress", will finish it this week and will shoot a video probably on weekend or next week. https://t.co/yZSE6NUSgU" / X
·x.com·
(20) Povilas Korop | Laravel & AI Coding Educator on X: "Tencent Hy3 model is surprisingly good for coding, for its price. Tested my new LLM benchmark projects on it, with OpenCode Go. Pleasantly surprised. P.S. v2 leaderboard is "work in progress", will finish it this week and will shoot a video probably on weekend or next week. https://t.co/yZSE6NUSgU" / X
(20) Povilas Korop | Laravel & AI Coding Educator on X: "Ok finished another round of LLM testing. Added Opus 5 High and GPT-5.6-Sol Medium to the table. Both ridiculously expensive compared to others. Best value for money is Luna medium/high. Full table here: https://t.co/qVlFywESzo https://t.co/NxFzHWOzy5" / X
(20) Povilas Korop | Laravel & AI Coding Educator on X: "Ok finished another round of LLM testing. Added Opus 5 High and GPT-5.6-Sol Medium to the table. Both ridiculously expensive compared to others. Best value for money is Luna medium/high. Full table here: https://t.co/qVlFywESzo https://t.co/NxFzHWOzy5" / X
·x.com·
(20) Povilas Korop | Laravel & AI Coding Educator on X: "Ok finished another round of LLM testing. Added Opus 5 High and GPT-5.6-Sol Medium to the table. Both ridiculously expensive compared to others. Best value for money is Luna medium/high. Full table here: https://t.co/qVlFywESzo https://t.co/NxFzHWOzy5" / X
(20) ollama on X: "Kimi K3 is now available on Ollama’s cloud. To use it with Claude Code, run: ollama launch claude --model kimi-k3:cloud Currently Kimi K3 requires a Pro or Max subscription, and consumes extra usage credits. We’re quickly working on adding capacity to expand access." / X
(20) ollama on X: "Kimi K3 is now available on Ollama’s cloud. To use it with Claude Code, run: ollama launch claude --model kimi-k3:cloud Currently Kimi K3 requires a Pro or Max subscription, and consumes extra usage credits. We’re quickly working on adding capacity to expand access." / X
·x.com·
(20) ollama on X: "Kimi K3 is now available on Ollama’s cloud. To use it with Claude Code, run: ollama launch claude --model kimi-k3:cloud Currently Kimi K3 requires a Pro or Max subscription, and consumes extra usage credits. We’re quickly working on adding capacity to expand access." / X
(21) Developers Criticize Anthropic's Claude Opus 5 as Downgrade from 4.8 / X
(21) Developers Criticize Anthropic's Claude Opus 5 as Downgrade from 4.8 / X
Launched on July 24, Opus 5 matches Opus 4.8's pricing at $5 per million input tokens and $25 per million output, while boasting benchmark wins like 79.2% on SWE-bench Pro and 44.4% on Frontier-Bench v0.1. Developers report it overcomplicates simple tasks, hallucinates details, makes more mistakes than 4.8, and requires heavy oversight in messy codebases. Power users like ThePrimeagen and Rod Johnson prefer rivals such as Fable 5 or GPT-5.6-sol, with many reverting to older models amid the hype-skepticism gap.
·x.com·
(21) Developers Criticize Anthropic's Claude Opus 5 as Downgrade from 4.8 / X
(21) Kimi.ai on X: "Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside https://t.co/Yz5uWeMbIm" / X
(21) Kimi.ai on X: "Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside https://t.co/Yz5uWeMbIm" / X
·x.com·
(21) Kimi.ai on X: "Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside https://t.co/Yz5uWeMbIm" / X
(20) Cursor on X: "Today we're launching Cursor Start, a new ₹649/month plan for developers in India. Start includes generous access to Grok 4.5 and Composer, so you can plan, build, test, and ship with agents every day. https://t.co/y2ZYlyzfR6" / X
(20) Cursor on X: "Today we're launching Cursor Start, a new ₹649/month plan for developers in India. Start includes generous access to Grok 4.5 and Composer, so you can plan, build, test, and ship with agents every day. https://t.co/y2ZYlyzfR6" / X
·x.com·
(20) Cursor on X: "Today we're launching Cursor Start, a new ₹649/month plan for developers in India. Start includes generous access to Grok 4.5 and Composer, so you can plan, build, test, and ship with agents every day. https://t.co/y2ZYlyzfR6" / X
(21) Kun Chen on X: "many people ask when to use a big model (like sol/fable) at low reasoning effort, vs a small model (like luna/sonnet) at high reasoning effort i deliberately forced myself to use all the permutations a lot over the last couple of weeks to build intuition, and i realized the https://t.co/0GtNczSVDE" / X
(21) Kun Chen on X: "many people ask when to use a big model (like sol/fable) at low reasoning effort, vs a small model (like luna/sonnet) at high reasoning effort i deliberately forced myself to use all the permutations a lot over the last couple of weeks to build intuition, and i realized the https://t.co/0GtNczSVDE" / X
·x.com·
(21) Kun Chen on X: "many people ask when to use a big model (like sol/fable) at low reasoning effort, vs a small model (like luna/sonnet) at high reasoning effort i deliberately forced myself to use all the permutations a lot over the last couple of weeks to build intuition, and i realized the https://t.co/0GtNczSVDE" / X
(10) Fireworks AI on X: "We ran Kimi K3 against Fable on ~1,000 agentic tasks, expecting a catch-up story. We got a specialization story instead. @kimi_moonshot's K3 outperformed on security, crypto, and long terminal loops. Fable beat on multi-lang + web/data viz. Per-task routing hits 93% accuracy, https://t.co/hvUAIYr4jZ" / X
(10) Fireworks AI on X: "We ran Kimi K3 against Fable on ~1,000 agentic tasks, expecting a catch-up story. We got a specialization story instead. @kimi_moonshot's K3 outperformed on security, crypto, and long terminal loops. Fable beat on multi-lang + web/data viz. Per-task routing hits 93% accuracy, https://t.co/hvUAIYr4jZ" / X
·x.com·
(10) Fireworks AI on X: "We ran Kimi K3 against Fable on ~1,000 agentic tasks, expecting a catch-up story. We got a specialization story instead. @kimi_moonshot's K3 outperformed on security, crypto, and long terminal loops. Fable beat on multi-lang + web/data viz. Per-task routing hits 93% accuracy, https://t.co/hvUAIYr4jZ" / X
asgeirtj/system_prompts_leaks: Extracted system prompts from Anthropic - Claude Fable 5, Opus 4.8, Claude Code, Claude Design. OpenAI - ChatGPT GPT-5.6, Codex GPT-5.6, GPT-5.5. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Cursor, Copilot, VS Code, Perplexity, and more. Updated regularly.
asgeirtj/system_prompts_leaks: Extracted system prompts from Anthropic - Claude Fable 5, Opus 4.8, Claude Code, Claude Design. OpenAI - ChatGPT GPT-5.6, Codex GPT-5.6, GPT-5.5. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Cursor, Copilot, VS Code, Perplexity, and more. Updated regularly.
·github.com·
asgeirtj/system_prompts_leaks: Extracted system prompts from Anthropic - Claude Fable 5, Opus 4.8, Claude Code, Claude Design. OpenAI - ChatGPT GPT-5.6, Codex GPT-5.6, GPT-5.5. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Cursor, Copilot, VS Code, Perplexity, and more. Updated regularly.