将图像堆叠为numpy数组比预分配更快吗？

Question

将图像堆叠为numpy数组比预分配更快吗？

4

我经常需要堆叠2D的numpy数组（tiff图像）。为此，我首先将它们附加在一个列表中，然后使用np.dstack。这似乎是获取3D阵列叠加图像最快的方法。但是，是否有更快/更节省内存的方法？

from time import time
import numpy as np

# Create 100 images of the same dimention 256x512 (8-bit). 
# In reality, each image comes from a different file
img = np.random.randint(0,255,(256, 512, 100))

t0 = time()
temp = []
for n in range(100):
    temp.append(img[:,:,n])
stacked = np.dstack(temp)
#stacked = np.array(temp)  # much slower 3.5 s for 100

print time()-t0  # 0.58 s for 100 frames
print stacked.shape

# dstack in each loop is slower
t0 = time()
temp = img[:,:,0]
for n in range(1, 100):
    temp = np.dstack((temp, img[:,:,n]))
print time()-t0  # 3.13 s for 100 frames
print temp.shape

# counter-intuitive but preallocation is slightly slower
stacked = np.empty((256, 512, 100))
t0 = time()
for n in range(100):
    stacked[:,:,n] = img[:,:,n]
print time()-t0  # 0.651 s for 100 frames
print stacked.shape

# (Edit) As in the accepted answer, re-arranging axis to mainly use 
# the first axis to access data improved the speed significantly.
img = np.random.randint(0,255,(100, 256, 512))

stacked = np.empty((100, 256, 512))
t0 = time()
for n in range(100):
    stacked[n,:,:] = img[n,:,:]
print time()-t0  # 0.08 s for 100 frames
print stacked.shape

- otterb

如果您保证temp中的所有数组都符合条件，则可以避免调用dstack。此时，您只需调用stacked = np.concatenate(temp,axis=2)来节省Python开销的小部分时间。如果您展示更多的代码，可能会有更好的解决方法，但是就所展示的代码而言，它已经几乎是最优的了。 - Daniel

Temp中的数组都是2D的，我想要连接它们以获得一个3D数组。因此，np.concatenate(temp, axis=2)会产生一个错误：轴2超出了范围[0, 2)。np.concatenate(temp, axis=1)将创建一个2D数组（256x51200）。 - otterb

我在评论中漏掉了一个关键部分，它应该是“...如果满足这个条件，temp 中的所有数组都是3D的”。需要注意的是，这种节省对于非常大的 temp 大小来说是微不足道的，可能每个数组约为2微秒。 - Daniel

1个回答

网页内容由stack overflow 提供, 点击上面的

可以查看英文原文，
原文链接

- Magellan88 · Accepted Answer

经过与otterb的共同努力，我们得出了预分配数组是正确的方式。显然，性能瓶颈是图像编号（n）作为最快变化索引的数组布局。如果我们将n作为数组的第一个索引（默认为“C”顺序：第一个索引变化最慢，最后一个索引变化最快），我们可以获得最佳性能：

from time import time
import numpy as np

# Create 100 images of the same dimention 256x512 (8-bit). 
# In reality, each image comes from a different file
img = np.random.randint(0,255,(100, 256, 512))

# counter-intuitive but preallocation is slightly slower
stacked = np.empty((100, 256, 512))
t0 = time()
for n in range(100):
    stacked[n] = img[n]
print time()-t0  
print stacked.shape