Python Tutorial
NumPy Data Types
Every ndarray has one dtype. Choosing int32 vs float64 changes memory use and what math you can do.
Check dtype
import numpy as np
print(np.array([1, 2, 3]).dtype) # int64 (platform dependent)
print(np.array([1.5, 2.0]).dtype) # float64
print(np.array(["a", "b"]).dtype) # <U1 (unicode)Set dtype When Creating
arr = np.array([1, 2, 3], dtype="i4") # 32-bit integer
print(arr.dtype)
print(arr)Common Codes
| Code | Meaning |
|---|---|
i / int32 | Signed integer |
u | Unsigned integer |
f / float64 | Float |
bool | Boolean |
S / U | Bytes / Unicode string |
M | Datetime64 |
Convert Types
arr = np.array([1.1, 2.9, 3.5])
ints = arr.astype("i")
print(ints) # [1 2 3] (truncated, not rounded)Casting a non-numeric string to int raises ValueError. Convert only when the data is actually numeric.
📘 Real-World Deep Dive
Knowing <strong>NumPy Data Types (NumPy)</strong> well is what turns NumPy from a curiosity into a daily tool — you'll reach for it in nearly every real project.
Real-Life Scenario
An end-to-end usage of NumPy Data Types that you'd actually see in a data pipeline or analytics notebook.
Real-Life Example
import numpy as np
a = np.array([1, 2, 3], dtype=np.int8)
b = a.astype(np.float32)
print(a.dtype, b.dtype)Expected Output
(see source)Common mistakes
- NumPy uses 0-based, C-order indexing — the rightmost axis is the *fastest-varying* one. Mixing it with Fortran-order arrays is a common surprise.
np.array([[1,2],[3,4]], dtype=int)is fine, but a ragged Python list produces dtype=object and silently disables vectorisation.- In-place ops (
a *= 2) sometimes break views instead of returning a new array; usenp.multiply(a, 2, out=...)if explicitness matters. - Treating NumPy Data Types as a black box without reading the docs — the API has subtle defaults that bite when you scale.
🚀 Performance & Best Practices
- Vectorise: replace Python
forloops with ufuncs; you can expect 10–100× speedups. - Pre-allocate output arrays with
np.emptyinstead of growing them withnp.append. - Keep data in float32 unless you need float64 precision — half the memory, double the cache locality.
- When working with NumPy, prefer vectorised / batched operations over Python loops.
🧪 Try It Yourself
- Reproduce the snippet on a representative slice of your own data.
- Profile the snippet with
cProfileortimeitand find the single biggest improvement. - Generalise the snippet into a small, reusable function you can drop into future projects.