Ceph pools define replication or erasure coding, placement groups organize objects for placement, and CRUSH rules select the OSDs that hold them. To configure a pool, distinguish replica count from capacity, inspect its PG autoscaler mode, and check that its placement rule matches the available hosts and devices. These settings interact: changing one can trigger data movement even when the application keeps using the same pool name.
What Ceph pools contain
An OSD, or object storage daemon, manages one storage device and is the unit of failure and capacity. Monitors maintain the authoritative cluster map and the consensus that keeps it consistent. Managers handle metrics and orchestration. A metadata server is added only when the file interface is used.
Data is written into pools. A pool is a logical namespace with its own durability scheme, its own placement rules, and its own placement group count. Block volumes, object buckets, and file data all end up as objects inside pools.
Ceph pool size: replica count, not terabytes
For a replicated pool, size counts all copies, including the primary. A value of three means three copies in total. Inspect the existing settings and usage first:
ceph osd pool ls detail
ceph osd pool get <pool> size
ceph osd pool get <pool> min_size
ceph df detail
Replace angle-bracket placeholders with an existing pool name. The pool operations reference documents this change to three replicas:
ceph osd pool set <pool> size 3
Use it only after checking that the selected failure domains and free capacity can hold three copies. Increasing size consumes more raw storage and requires replication work; it does not increase usable capacity. min_size separately controls how many replicas must be available for I/O. Reducing either setting to clear a warning can sacrifice durability.
An erasure-coded pool reports size=k+m, but that value cannot be set directly. Changing k or m requires a new pool and data migration. For space estimates, use the redundancy profile rather than interpreting size as a capacity limit.
Placement groups sit between objects and devices
Ceph does not map each object directly to devices. Objects hash into placement groups, and placement groups map onto sets of OSDs. That indirection is what keeps the mapping cheap to compute and recovery manageable, since the cluster tracks and repairs a bounded number of placement groups rather than an unbounded number of objects.
Placement group count is a real tuning decision. Too few and data distributes unevenly, leaving some OSDs much fuller than others while capacity sits idle elsewhere. Too many and monitors and OSDs carry more peering state and memory overhead than they need to. Ceph includes an autoscaler that adjusts counts based on how much data a pool actually holds relative to the cluster, and using it is usually better than a hand picked number that stops being right as the cluster grows. Both extremes surface as named health checks, POOL_TOO_FEW_PGS and TOO_MANY_PGS, and what to do about each is set out in diagnosing and clearing Ceph HEALTH_WARN.
Inspect PG autoscaler behavior before changing counts
The placement group documentation defines three per-pool modes:
| Mode | Behavior |
|---|---|
off | The operator chooses the PG count. |
warn | Ceph recommends adjustments through health checks. |
on | Ceph adjusts the PG count automatically. |
ceph osd pool autoscale-status
ceph osd pool get <pool> pg_autoscale_mode
Read the current PG count, any proposed new count and the mode together. The recommendation accounts for pool usage and redundancy overhead. A blank recommendation does not mean the pool has no PGs; the autoscaler avoids proposing every small change in its estimate.
After reviewing the recommendation, enable automatic adjustment for that pool with:
ceph osd pool set <pool> pg_autoscale_mode on
Allow time for the adjustment and associated movement, then inspect health again. Do not copy a fixed PG count from a differently sized cluster. Capacity additions can change the appropriate count, so repeat this review after expansion.
CRUSH turns a name into a location
CRUSH is the algorithm that maps a placement group onto specific OSDs. It takes the cluster topology, expressed as a hierarchy of devices inside hosts inside racks inside rooms, plus a set of rules, and deterministically computes which OSDs should hold a given placement group. Every client runs the same algorithm against the same map and gets the same answer, so there is no lookup service in the data path and no metadata bottleneck.
For a comparison with another distributed storage system, see how Ceph and Gluster place data.
Because CRUSH is computed rather than looked up, the shape of the physical cluster is an input to correctness and not merely to performance. How many hosts, how they are racked, and what each one is built from is the first decision rather than the last; the published figures for OSD processors, memory, drives and networking are collected in Ceph hardware requirements.
The most consequential thing in a CRUSH rule is the failure domain. It states the level of the hierarchy across which copies must be separated. With host as the failure domain, no two copies of the same data land on the same server, so losing a server never loses all copies. With rack as the failure domain, copies are spread across racks and a rack level power or switch failure is survivable. The failure domain must be matched by real capacity: a rule demanding three copies in distinct racks cannot be satisfied by two racks, and the affected placement groups will simply stay undersized.
Device classes let one cluster hold mixed media and route pools to the right tier, with rules selecting only flash or only spinning devices.
Change a Ceph pool’s CRUSH rule
The CRUSH map reference describes rules by pool type, failure domain and optional device class. Inspect the current assignment and the proposed rule before making a change:
ceph osd pool get <pool> crush_rule
ceph osd crush rule ls
ceph osd crush rule dump <rule-name>
ceph osd tree
Record the current rule name. Confirm the proposed rule supports the pool type and selects enough distinct hosts or racks with sufficient free space. A flash-only rule cannot use new HDD OSDs. For an existing EC pool, keep the coding profile compatible; changing the rule does not change k or m.
Apply the existing, reviewed rule to the named pool:
ceph osd pool set <pool> crush_rule <rule-name>
ceph osd pool get <pool> crush_rule
ceph status
ceph health detail
Reassignment can trigger substantial backfill. Check PG states and per-OSD fullness while it completes. If the rule needs more eligible capacity, first add an OSD to Ceph on an appropriate host; another disk on the same host does not add a host failure domain.
Replication or erasure coding
Replicated pools store whole copies. Writes go to a primary OSD which forwards to the others, and reads are served locally. This costs the most raw capacity per usable byte and recovers quickly, because recovery is a copy.
Erasure coded pools split an object into data chunks plus computed parity chunks, spread across more OSDs. Usable capacity per raw byte is much better, but every write and every degraded read involves computation and touches more devices, so latency and CPU cost rise and recovery is heavier. Erasure coding suits large sequential workloads such as object storage and backups. Small random writes are the worst case for it.
The choice is worth making with numbers rather than instinct, because the profile cannot be altered after a pool is created. Erasure coding versus replication sets out the space amplification of each common profile, how many failure domains each one demands, and where the write penalty actually lands. The Ceph erasure coding calculator estimates raw and theoretical usable capacity plus the sum of the default OSD memory targets.
Common mistakes
Setting a failure domain the physical layout cannot satisfy and leaving placement groups permanently degraded. Treating near full OSDs as a warning to be dismissed, when they block writes and stall recovery. Building a cluster with too few failure domains to tolerate the loss of one. Putting latency sensitive block workloads on erasure coded pools. Ignoring uneven OSD utilisation instead of investigating placement group counts and device weights.