🥒7 Reasons Not to Use Pickle to Save ML Models
A Data Scientist often writes code in notebooks like Jupyter Notebook, Google Colab, or specialized IDEs. To port this code to a production environment, it must be converted to a lightweight interchange format, compressed and serialized, that is independent of the development language. One of these formats is Pickle, a binary version of a Python object for serializing and deserializing its structure, converting a hierarchy of Python objects into a stream of bytes and vice versa. The Pickle format is quite popular due to its lightness. It does not require a data schema and is quite common, but it has a number of disadvantages:
• Unsafe. You can only unpack pickle files that you trust. An attacker can create malicious data that will execute arbitrary code during decompression. You can mitigate this risk by signing the data with hmac to make sure it hasn't been tampered with. The unreliability comes not from the fact that Pickles contain code, but from the fact that they create objects by calling the constructors mentioned in the file. Any callable object can be used instead of a class name to create objects. Malicious code will use other Python callables as constructors.
• Code mismatch. If the code changes between the time the ML model is packaged in the Pickle file and when it is used, the objects may not match the code. They will still have the structure created by the old code, but will try to work with the new version. For example, if an attribute was added after the Pickle was created, the objects in the Pickle file will not have that attribute. And if the new version of the code is supposed to process it, there will be problems.
• Implicit serialization. On the one hand, the Pickle format is convenient in that it serializes any structure of a Python object. But at the same time, there is no way to specify preferences for serialization of one or another type of data. Pickle serializes everything in objects, even data that doesn't need to be serialized. But there is no way to skip the serialization of this or that attribute in Pickle. If an object contains an attribute that cannot be boxed, such as an object with an open file, Pickle will not skip it, insisting on trying to box it, and then throw an exception.
• Lack of initialization. Pickle stores the entire structure of objects. When the Pickle module recreates the objects, it does not call the init method because the object has already been created, considering the initialization to have been called when the object was first created during the creation of the Pickle file. But the init method can do some important things, like opening file objects. In this case, raw objects will be in a state that is incompatible with the init method. Or initialization can register information about the object being created. Then unselected objects will not be displayed in the general log.
• Unreadable. Pickle is a stream of binary data, i.e. instructions for the abstract execution mechanism. Once a Pickle is opened as a normal file, its contents cannot be read. To find out what is in it, you will have to use the Pickle module to load. This can make debugging difficult, since it is difficult to find the desired data in binaries.
• Binding to Python. Being a Python library, Pickle is specific to this programming language. Although the format itself can be used for other programming languages, it is difficult to find packages that provide such capabilities. Also, they will be limited to cross-language common list/dict object structures. Pickle serializes objects containing callable functions and classes without problems. But the format does not store the code, but only the name of the function or class. When unpacking data, function names are used to look for existing code in the running process.
• Low speed. Finally, compared to other serialization methods, Pickle is much slower.
Post #488
768
- 👍 3