AI models that are open

AI models that are open 


Background

  AI models that are "open" tend to be open-weight models (with model parameters (weights and biases) and the inference code to run it) that you can download to your local environment and run it, instead of OSAID compliant open-source.  Open AI models are popular due to the ability to "fine tune" (customize or adapt) them to your specific purpose, rather than being more general purpose.  Fine tuning a model is the process of setting some new hyperparameter values and giving a trained model a new, specialized dataset and retrain it so it creates a new model with new weights.  (Usually followed by "model merges" (combines the two models together) and shrinks the model (by compression (like GGUF, AWQ, or EXL2) to reduce the precision of the weights (e.g., converting 16-bit numbers into 4-bit numbers)) for use on home computers. 


Executive Summary

   America's main open models are Meta's Llama and Google's Gemma. Chinese open models are DeepSeek, Alibaba's Qwen, Moonshot AI (Kimi), Zhipu AI, and Tencent.  Cumulative downloads show that China now has the lead in downloads of its open AI models. Alibaba's Qwen has overtaken Meta's Llama in downloads. Then after downloading a open model, the person fine tunes the model and posts it to HuggingFace ( https://huggingface.co/ ). A good one I like is ThinkingCap-Qwen3.6-27B  ( https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B ) with local version at ( https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF ).  Another is ( https://huggingface.co/GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF ).




Source: https://atomproject.ai/





Source: https://atomproject.ai/


Historical Tech
   So far, the big historical improvements in AI have been Google's "The Transformer" (attention mechanism) and "Inference-Time Scaling" (AI can spend more compute resources and time on inference to produce better results so small ones run in big compute environments can do well). Future improvements are "Reinforcement Learning with Verifiable Rewards" (RLVR) and distillation (where a large, powerful model tutors a smaller one).

What Next
  The software company Cursor, used Moonshot AI’s Kimi K2.5 as the basis for its most recent family of models specialized for coding.

Comments

Popular posts from this blog

GHL Email Campaigns

Await

Free AI Tools