DeepSeek Opens V4.1 Flash Beta With Native Multimodality and Higher Speed
Summary
DeepSeek has opened an internal test of an intermediate V4.1 Flash version, identified as deepseek-v4.1-flash-expires-on-0910 and scheduled to expire on September 10. The company says the model uses a new architecture with native multimodal input and output, but it has not released a technical report or performance table. Community tests cited by the article report a 0.3-second thinking time, 159.3 tokens per second, and 0.8 seconds for an end-to-end exchange in one short prompt; another long-text reasoning test reached 420 tokens per second and 409.5 tokens per second end to end. Compared with V4 Flash Vision-Exp, testers reported speedups of 5.2 times for 49k-context retrieval, 6.0 times for SVG generation, 4.6 times for a Manacher algorithm problem, 5.0 times for large SQL generation and optimization, and 3.9 times for an asyncio refactoring task. A photo test also suggested that the model correctly identified a striped suit, addressing a multimodal failure the tester had initially suspected. The article says V4.1 Flash retains V4 Flash’s calling price. Around the same time, DeepSeek announced 150 engineering openings covering model research infrastructure, Agent frameworks, APIs, online services, data engineering, and elastic computing. A company executive attributed the hiring to rapidly increasing data, machines, training and evaluation tasks, Agent environments, users, and requests, which require more scalable scheduling, inference, and backend systems.