Running open-weight models locally or on your own servers keeps data in-house, but shifts security responsibility to you.
Safe Model Files
Some model file formats can execute arbitrary code when loaded. Prefer safe formats such as safetensors or GGUF, download from verified publishers, and check hashes.
Serving Software
Local inference servers are software like any other. Keep them updated, and don't expose their APIs to networks without authentication — many default to no authentication at all.
Access Control
Put an authenticating gateway in front of model endpoints, with rate limits and logging.
Safety Behaviour
Open-weight models may have weaker safety training, or modified versions may have had it removed. Add your own input and output safeguards for user-facing applications.
Data Handling
Local models avoid sending data to providers, but prompts and outputs may still be logged locally. Apply retention and access policies.
Resource Protection
GPU servers are attractive to attackers for cryptocurrency mining and other misuse. Monitor usage.
Licences
Check model licences for restrictions on use.