DeepSeek's V4 Flash model struggles with tasks, price increase planned
DeepSeek's V4 Flash model has only completed 53.8% of complex tasks in recent evaluations, revealing significant performance issues despite its high ranking. As a result, DeepSeek plans to raise pricโฆ
DeepSeek's V4 Flash model, which has gained significant acclaim since its launch, has struggled in practical applications, completing only 53.8% of complex agent tasks in recent tests. The evaluation, conducted by Composio, involved running V4 Flash through eight different agent harnesses, including well-known tools like Claude Code, Codex, and OpenCode. The model faced 30 intentionally challenging multi-step tasks that included interactions with popular platforms such as Gmail, GitHub, Slack, and Google Sheets.
The testing results highlight a critical issue: while V4 Flash has achieved a top rank in model leaderboards, its performance varies significantly based on the harness and configuration used. Out of 240 total runs, only 129 were successful, and just six workflows were completed without any failures across all tested harnesses. This disparity suggests that effective orchestration is more crucial than raw model capabilities when it comes to deploying AI in real-world enterprise scenarios. Factors like tool configuration, caching behavior, retries, and the provider stack played a significant role in the varying outcomes.
In light of these mixed results, DeepSeek plans to increase the prices for both V4 Flash and its Pro models. These models have quickly become popular choices among developers aiming to create coding assistants and other AI-driven agents. The price hike indicates confidence in their long-term value, despite the current performance issues highlighted by the testing.
The implications of these findings are significant for businesses considering the integration of AI models like V4 Flash. Organizations will need to weigh the model's capabilities against its real-world performance and the importance of orchestration in achieving desired outcomes. As companies navigate the challenges of deploying AI in complex environments, the effectiveness of these models may ultimately determine their success in the market.
Read Full Story at VentureBeat โ


