DIRD An Introduction to the Statistical Drake Equation
DIA / AAWSAP contractor
Public domain · full text
In one page
This is the Defense Intelligence Reference Document about arithmetic rather than hardware. Its author — Appendix B reproduces Claudio Maccone’s 2008 congress paper under his own name — takes Frank Drake’s 1961 equation for the number of communicating civilisations in the galaxy and points out that multiplying seven guesses together gives one tidy number that hides how uncertain every guess was. His repair is to replace each of the seven factors with a random variable carrying a mean value and a spread. Taking logarithms turns the product into a sum, and the central limit theorem then forces the answer into a lognormal distribution whose mean, spread, median and peak can all be written down exactly. Worked through with one plausible set of inputs, the classical equation returns 3,500 civilisations; the statistical version returns a mean of about 4,590, and makes the distance to the nearest one a random variable too — most likely 1,933 light years, with a 75 percent chance of lying between 1,361 and 3,979 light years.
Why it matters hereChapter 1’s point about the DIRD list is that the Defense Intelligence Agency commissioned an engineering and physics reading list, and this is the page of it that teaches the manners the rest of the list is read with: state the spread, not just the number, and let the error bars do the arguing. Chapter 13’s unified picture needs exactly that discipline whenever the subject turns to who else is out there and how far away they are.
What it claims
01The classical Drake equation is over-simplified: it multiplies seven uncertain numbers into a single sheer integer, when each factor should instead be a random variable carrying both a mean value and a standard deviation.Section 4, p. 10; Section 5, p. 11
Published and peer-reviewed02Because each Drake factor is non-negative and its true distribution is unknown, Shannon’s 1948 theorem — that the uniform distribution carries maximum entropy over any finite range — makes the uniform distribution the honest choice for every input.Section 5, p. 12; Appendix A, pp. 25–27
Settled physics03Taking natural logarithms turns the product into a sum, so the central limit theorem makes the number of communicating civilisations N a lognormal random variable whose mean, variance, mode, median and every moment follow analytically.Section 6, p. 13; Table 1, p. 14
Published and peer-reviewed04For one worked set of inputs the classical equation returns N = 3500, while the statistical version returns a mean of about 4590 with a standard deviation of 11,195 — the statistical extension raises rather than lowers the expected number.Section 7, pp. 16–17
Published and peer-reviewed05The distance to the nearest civilisation becomes a random variable with its own probability density: mode 1,933 light years, mean 2,670 light years, standard deviation 1,309 light years, and a 75 percent probability of lying between 1,361 and 3,979 light years.Section 8, pp. 18–22, equations (9) through (15)
Published and peer-reviewed06The Data Enrichment Principle: because the central limit theorem allows any number of factors, new factors can be added to the equation as scientific knowledge grows, until it becomes a growing computer code — humanity’s first Encyclopaedia Galactica.Section 9, p. 23; Section 10: Conclusions, p. 23
What to watch
Read it
An Introduction to the Statistical Drake Equation
Defense Intelligence Reference Document, Acquisition Threat Support. 11 March 2010 (IOD: 1 December 2009).
Prepared by the Defense Intelligence Agency. This product is one in a series of advanced technology reports produced in FY 2009 under the Defense Intelligence Agency Advanced Aerospace Weapon System Applications (AAWSA) Program.
1. Introduction
SETI (an acronym for "Search for Extraterrestrial Intelligence") is a relatively new branch of scientific research, having begun only in 1959. Its goal is to ascertain whether alien civilizations exist in the universe, how far from us they exist, and possibly how much more advanced than us they may be.
As of 2009, the only physical tools we know that could help us get in touch with aliens are the electromagnetic waves an alien civilization could emit and we could detect. This forces us to use the largest radiotelescopes on Earth for SETI research, because the higher our collecting area of electromagnetic radiation is, the higher our sensitivity is (that is, the farther in space we can probe). Yet, even by using the largest radiotelescopes on Earth (the 310-meter dish at Arecibo, for instance), we cannot search for aliens beyond, say, a few hundred light years away. This is a very, very small amount of space around us within our galaxy, the Milky Way, that is about 100,000 light years in diameter. Thus, current SETI can cover only a very tiny fraction of the galaxy, and it is not surprising that in the past 50 years of SETI searches, NO extraterrestrial civilization was discovered. Quite simply, we did not get far enough!
This demands the construction of much more powerful and radically new radiotelescopes. Rather than big and heavy metal dishes, whose mechanical problems hamper SETI research too much, we are now turning to "software radiotelescopes," where a large number of small dishes (ATA = Allen Telescope Array, and ALMA = Atacama Large Millimeter/submillimeter Array) or even just of simple dipoles (LOFAR = Low Frequency Array) using state-of-the-art electronics and very-high-speed computing can outperform the classical radiotelescopes in many regards. The final dream in this field is the SKA (= Square Kilometer Array), currently being designed and expected to be completed around 2020.
2. The Key Question: How Far are They?
But still, the key question remains: how far are they?
Or, more correctly, how far do we expect the NEAREST extraterrestrial civilization to be from the Solar System in the galaxy?
This question was first faced in a scientific manner back in 1961 by the same scientist who also was the first experimental SETI radio astronomer ever: the American, Frank Donald Drake (born 1930). He first considered the shape and size of the galaxy where we are living: the Milky Way. This is a spiral galaxy measuring some 100,000 light years in diameter and some 16,000 light years in thickness of the Galactic Disk at half-way from its center. That is:
The diameter of the galaxy is (about) 100,000 light years, (abbreviated ly) i.e., its radius, R_Galaxy, is about 50,000 ly.
The thickness of the Galactic Disk at half-way from its center, h_Galaxy, is about 16,000 ly.
The volume of the galaxy may then be approximated as the volume of the corresponding cylinder, i.e.
V_galaxy = pi · R_Galaxy² · h_Galaxy (1)
Now consider the sphere around us having a radius r. The volume of such a sphere is
V_our_sphere = (4/3) · pi · (ET_Distance / 2)³ (2)
In the last equation, we had to divide the distance "ET_Distance" between ourselves and the nearest ET civilization by 2 because we are now going to make the unwarranted assumption that all ET civilizations are equally spaced from each other in the galaxy! This is a crazy assumption, clearly, and should be replaced by more scientifically-grounded assumptions as soon as we know more about our Galactic Neighborhood. At the moment, however, this is the best guess that we can make, and so we shall take it for granted, although we are aware that this is a weak point in the reasoning.
Furthermore, let us denote by N the total number of civilizations now living in the galaxy, including ourselves. Of course, this number N is unknown. We only know that N is at least 1 since one civilization does at least exist!
Having thus assumed that ET civilizations are UNIFORMLY SPACED IN THE GALAXY, we can then write down the proportion:
V_galaxy / N = V_our_sphere (3)
That is, upon replacing both (1) and (2) into (3):
pi · R_Galaxy² · h_Galaxy / N = (4/3) · pi · (ET_Distance / 2)³ (4)
The last equation contains two unknowns: N and ET_Distance, and so we don't know which one it is better to solve for.
However, we may suppose that, by resorting to the (rather uncertain) knowledge that we have about the Evolution of the galaxy through the last 10 billion years or so, we might somehow compute an approximate value for N.
Then, we may solve (4) for ET_Distance thus obtaining the (AVERAGE) DISTANCE BETWEEN ANY PAIR OF NEIGHBORING CIVILIZATIONS IN THE GALAXY (DISTANCE LAW)
ET_Distance(N) = cube root of (6 · R_Galaxy² · h_Galaxy) / cube root of N = C / cube root of N (5)
where the positive constant C is defined by
C = cube root of (6 · R_Galaxy² · h_Galaxy) ≈ 28,845 light years (6)
Equations (5) and (6) are the starting point to understand the origin of the Drake equation that we discuss in detail in Section 3 of this paper.
Let us just complete this section by pointing out three different numerical cases of the distance law (5):
-
We know that we exist, so N may not be smaller than 1, i.e., N is at least 1. Suppose then that we are alone in the galaxy, i.e., that N = 1. Then the distance law (5) yields as distance to the nearest civilization from us just the constant C, i.e. 28,845 light years. This is about the distance in between ourselves and the center of the galaxy (i.e. the Galactic Bulge). Thus, this result seems to suggest that, if we do not find any extraterrestrial civilization around us in these outskirts of the galaxy where we live, we should look around the Galactic Center first. And this is indeed what is happening, i.e., many SETI searches are actually pointing the antennas towards the Galactic Center, looking for beacons.
-
Suppose next that N = 1000, i.e. there are about a thousand extraterrestrial communicating civilizations in the whole galaxy right now. Then the distance law (5) yields an average distance of 2,885 light years. This is a distance that most radiotelescopes on Earth may not reach for SETI searches right now: hence the need to build larger radiotelescopes, like ALMA, LOFAR and the SKA.
-
Suppose finally that N = 1,000,000, i.e., there are a million communicating civilizations now in the galaxy. Then the distance law (5) yields an average distance of 288 light years. This is within the (upper) range of distances that our current radiotelescopes may reach for SETI searches, and that justifies all SETI searches that have been done so far in the first fifty years of SETI (1960–2010).
In conclusion, interpolating the above three special cases of N, we may say that the distance law (5) yields the following key diagram of the average ET distance vs. the assumed number of communicating civilizations, N, in the galaxy right now (Figure 1):
Figure 1. DISTANCE LAW; i.e., the Average Distance (plot along the vertical axis in light years) Versus the NUMBER of Communicating Civilizations ASSUMED to Exist in the Galaxy Right Now.
3. Computing N By Virtue of the Drake Equation (1961)
In the previous section, the problem of finding how close the nearest ET civilization may be was "solved" by reducing it to the computation of N, the total number of extraterrestrial civilizations now existing in this galaxy. In this section the famous Drake equation is described, that was proposed back in 1961 by Frank Donald Drake (born 1930) to estimate the numerical value of N. We believe that no better introductory description of the Drake equations exists other than the one given by Carl Sagan in his 1983 book "Cosmos," in its turn based on the famous TV series "Cosmos." So, in this paragraph we report Carl Sagan's description of the Drake equation unabridged.
"But is there anyone out there to talk to? With a third or a half a trillion stars in our Milky Way galaxy alone, could ours be the only one accompanied by an inhabited planet? How much more likely it is that technical civilizations are a cosmic commonplace, that the galaxy is pulsing and humming with advanced societies, and, therefore, that the nearest such culture is not so very far away — perhaps transmitting from antennas established on a planet of a naked-eye star just next door. Perhaps when we look up at the sky at night, near one of those faint pinpoints of light is a world on which someone quite different from us is then glancing idly at a star we call the Sun and entertaining, for just a moment, an outrageous speculation.
It is very hard to be sure. There may be several impediments to the evolution of a technical civilization. Planets may be rarer than we think. Perhaps the origin of life is not so easy as our laboratory experiments suggest. Perhaps the evolution of advanced life forms is improbable. Or it may be that complex life forms evolve more readily, but intelligence and technical societies require an unlikely set of coincidences — just as the evolution of the human species depended on the demise of the dinosaurs and the ice-age recession of the forests in whose trees our ancestors screeched and dimly wondered. Or perhaps civilizations arise repeatedly, inexorably, on innumerable planets in the Milky Way, but are generally unstable; so all but a tiny fraction are unable to survive their technology and succumb to greed and ignorance, pollution and nuclear war.
It is possible to explore this great issue further and make a crude estimate of N, the number of advanced civilizations in the galaxy. We define an advanced civilization as one capable of radio astronomy. This is, of course, a parochial if essential definition. There may be countless worlds on which the inhabitants are accomplished linguists or superb poets but indifferent radio astronomers. We will not hear from them. N can be written as the product or multiplication of a number of factors, each a kind of filter, every one of which must be sizable for there to be a large number of civilizations:
- Ns, the number of stars in the Milky Way galaxy.
- fp, the fraction of stars that have planetary systems.
- ne, the number of planets in a given system that are ecologically suitable for life.
- fl, the fraction of otherwise suitable planets on which life actually arises.
- fi, the fraction of inhabited planets on which an intelligent form of life evolves.
- fc, the fraction of planets inhabited by intelligent beings on which a communicative technical civilization develops.
- fL, the fraction of planetary lifetime graced by a technical civilization.
Written out, the equation reads
N = Ns · fp · ne · fl · fi · fc · fL (7)
All of the f's are fractions, having values between 0 and 1; they will pare down the large value of Ns."
4. The Drake Equation is Over-Simplified
In the nearly fifty years (1961–2009) elapsed since Frank Drake proposed his equation, a number of scientists and writers tried to find out which numerical values of its seven independent variables are more realistic in agreement with our present-day knowledge. Thus there is a considerable amount of literature about the Drake equation nowadays, and, as one can easily imagine, the results obtained by the various authors largely differ from one another. In other words, the value of N, that various authors obtained by different assumptions about the astronomy, the biology and the sociology implied by the Drake equation, may range from a few tens (in the pessimist's view) to some million or even billions in the optimist's opinion. A lot of uncertainty is thus affecting our knowledge of N as of 2010. In all cases, however, the final result about N has always been a sheer number, i.e., a positive integer number ranging from 1 to millions or billions. This is precisely the aspect of the Drake equation that this author regarded as "too simplistic" and improved mathematically in his paper #IAC-08-A4.1.4, entitled "The Statistical Drake Equation" and presented on October 1st, 2008, at the 59th International Astronautical Congress (IAC) held in Glasgow, Scotland, UK (September 29th thru October 3rd, 2008). That paper is attached herewith as Appendix B. Newcomers to SETI and to the Drake equation, however, may find that paper too difficult to be understood mathematically at a first reading. Thus, I shall now explain the content of that paper "by speaking easily." I thank the reader for his or her attention.
5. The Statistical Drake Equation
We start by an example.
Consider the first independent variable in the Drake equation (7), i.e., Ns, the number of stars in the Milky Way galaxy. Astronomers tell us that approximately there should be about 350 millions stars in the galaxy. Of course, nobody has counted (or even seen in the photographic plates) all the stars in the galaxy! There are too many practical difficulties preventing us from doing so: just to name one, the dust clouds that don't allow us to see even the Galactic Bulge (i.e. the central region of the galaxy) in the visible light (although we may "see it" at radio frequencies like the famous neutral hydrogen line at 1420 MHz). So, it doesn't make any sense to say that Ns = 350 × 10⁶, or, say (even worse) that the number of stars in the galaxy is (say) 354,233,321, or similar fanciful exact integer numbers. That is just silly and non-scientific. Much more scientific, on the contrary, is to say that the number of stars in the galaxy is 350 million plus or minus, say, 50 millions (or whatever values the astronomers may regard as more appropriate, since this is just an example to let the reader understand the difficulty).
Thus, it makes sense to REPLACE each of the seven independent variables in the Drake equation (7) by a MEAN VALUE (350 millions, in the above example) PLUS OR MINUS A CERTAIN STANDARD DEVIATION (50 millions, in the above example).
By doing so, we have made a great step ahead: we have abandoned the too-simplistic equation (7) and replaced it by something more sophisticated and scientifically more serious: the STATISTICAL Drake equation. In other words, we have transformed the classical and simplistic Drake equation (7) into an advanced statistical tool for the investigation of a host of facts hardly known to us in detail. In other words still:
- We replace each independent variable in (7) by a RANDOM VARIABLE, labeled Di (from Drake).
- We assume that the MEAN VALUE of each Di is the same numerical value previously attributed to the corresponding independent variable in (7).
- But now we also ADD A STANDARD DEVIATION sigma_Di on each side of the mean value, that is provided by the knowledge gathered by scientists in each discipline encompassed by each Di.
Having so done, the next question is:
How can we find out the PROBABILITY DISTRIBUTION for each Di?
For instance, shall that be a Gaussian, or what?
This is a difficult question, for nobody knows, for instance, the probability distribution of the number of stars in the galaxy, not to mention the probability distribution of the other six variables in the Drake equation (7).
There is a brilliant way to get around this difficulty, though.
We start by excluding the Gaussian because each variable in the Drake equation is a POSITIVE (or, more precisely, a non-negative) random variable, while the Gaussian applies to REAL random variables only. So, the Gaussian is out. Then, one might consider the large class of well-studied and positive probability densities called "the gamma distributions," but it is then unclear why one should adopt the gamma distributions and not any other. The solution to this apparent conundrum comes from Shannon's Information Theory and a theorem that he proved in 1948: "The probability distribution having maximum entropy (= uncertainty) over any FINITE range of real values is the UNIFORM distribution over that range." This is proven in Appendix A of the present document.
So, at this point, we assume that each of the seven Di in (7) is a UNIFORM random variable, whose mean value and standard deviation is known by the scientists working in the respective field (let it be astronomy, or biology, or sociology). Notice that, for such a uniform distribution, the knowledge of the mean value mu_Di and of the standard deviation sigma_Di automatically determines the RANGE of that random variable in between its lower (called ai) and upper (called bi) limits: in fact these limits are given by the equations
ai = mu_Di − sqrt(3) · sigma_Di bi = mu_Di + sqrt(3) · sigma_Di (8)
(the "surprising" factor sqrt(3) in the above equations comes from the definitions of mean value and standard deviation: please see equations (12), (15) and (17) in Appendix B for the relevant proof). So the uniform distribution of each random variable Di is perfectly determined by its mean value and standard deviation, and so are all its other properties.
The next problem is the following:
OK, since we now know everything about each uniformly distributed Di, what is the probability distribution of N, given that N is the product (7) of all the Di?
In other words, not only do we want to find the analytical expression of the probability density function of N, but we also want to relate its mean value mu_N to all mean values mu_Di of the Di, and its standard deviation sigma_N to all standard deviations sigma_Di of the Di.
This is a difficult problem.
It occupied the author's mind for no less than about ten years (1997–2007). It is actually an ANALYTICALLY UNSOLVABLE problem, in that, to the best of this author's knowledge, it is IMPOSSIBLE to find an analytic expression for any FINITE PRODUCT of uniform random variables Di. This result is proven in Sections 2 thru 3.3 of Appendix B (unfortunately!).
6. Solving the Statistical Drake Equation By Virtue of the Central Limit Theorem (CLT) of Statistics
The solution to the problem of finding the analytical expression for the probability density function of N in the statistical Drake equation was found by this author in September 2007. The key steps are the following:
- Take the natural logs of both sides of the statistical Drake equation (7). This changes the product into a sum.
- The mean values and standard deviations of the logs of the random variables Di may all be expressed analytically in terms of the mean values and standard deviations of the Di.
- Recall the Central Limit Theorem (CLT) of statistics, stating that (loosely speaking) if you have a SUM of independent random variables, each of which is ARBITRARILY DISTRIBUTED (hence, also including uniformly distributed), then, when the number of terms in the sum increases indefinitely (i.e. for a sum of random variables infinitely long) … the SUM RANDOM VARIABLE TENDS TO A GAUSSIAN.
- Thus, the natural log of N tends to a Gaussian.
- Thus, N tends to the LOGNORMAL DISTRIBUTION.
- The mean value and standard deviations of this lognormal distribution of N may all be expressed analytically in terms of the mean values and standard deviations of the logs of the Di already found previously.
This result is fundamental.
All the relevant equations are summarized in the following Table 1. This table is actually the same as Table 2 of the author's original paper IAC-08-A4.1.4, entitled "The Statistical Drake Equation" and presented by him at the International Astronautical Congress (IAC) held in Glasgow, UK, on October 1st, 2008. This original paper is reproduced in Appendix B.
To sum up, not only is it found that N approaches the completely known lognormal distribution for an INFINITY of factors in the statistical Drake equation (7), but the way is paved to further applications by removing the condition that the number of terms in the product (7) must be FINITE.
This possibility of ADDING ANY NUMBER OF FACTORS IN THE DRAKE EQUATION (7) was not envisaged, of course, by Frank Drake back in 1961, when "summarizing" the evolution of life in the galaxy in SEVEN simple STEPS. But today, the number of factors in the Drake equation should already be increased: for instance, there is no mention in the original Drake equation of the possibility that asteroidal impacts might destroy the life on Earth at any time, and this is because the demise of the dinosaurs at the K/T impact had not been yet understood by scientists in 1961, and was so only in 1980!
In practice, the number of factors should INCREASE as much as necessary in order to get better and better estimates of N as long as our scientific knowledge increases. This is called the "Data Enrichment Principle" and I believe it should be the next important goal in the study of the statistical Drake equation.
Finally, a numerical example explaining how the statistical Drake equation works in the practice will be given in the next section.
Table 1. Summary of the Properties of the Lognormal Distribution That Applies to the Random Variable N = Number of ET Communicating Civilizations in the Galaxy. The table gives, in terms of the two lognormal parameters mu and sigma, the probability density function, the mean value e^mu · e^(sigma²/2), the variance, the standard deviation, all the k-th moments, the mode e^mu · e^(−sigma²), the peak value of the density, the median e^mu, the skewness and the kurtosis; and it expresses mu and sigma² directly in terms of the lower (ai) and upper (bi) limits of the seven uniform Drake input random variables Di. (The formulae in this table did not survive the released scan legibly and are given in full in the source.)
7. An Example Explaining the Statistical Drake Equation
To understand how things work in practice for the statistical Drake equation, please consider the following Table 2. It is made up of three columns:
- The first column on the left lists the seven input sheer numbers that also become
- The mean values (middle column).
- Finally the last column on the right lists the seven input standard deviations.
The bottom line is the classical Drake equation (7). We see that, for this particular set of seven inputs, the classical Drake equation (i.e. the product of the seven numbers) yields a total of 3500 communicating extraterrestrial civilizations existing in the galaxy right now.
Table 2. Input Values (i.e. mean values and standard deviations) for the Seven Drake Uniform Random Variables Di. The first column on the left lists the seven input sheer numbers that also become the mean values (middle column). Finally the last column on the right lists the seven input standard deviations. The bottom line is the classical Drake equation (7). (The seven numerical input pairs did not survive the released scan legibly; the document states the bottom line as N = 3500.)
The statistical Drake equation, however, provides a much more articulated answer than just the above sheer number N = 3500. In fact, a MathCad code written by this author and capable of performing all the numerical calculations required by the statistical Drake equation for a given set of seven input mean values plus seven input standard deviations, yields for N the lognormal distribution (thin curve) plotted in Figure 2. We see immediately that the peak of this thin curve (i.e. the mode) falls at about n_mode = n_peak = e^mu · e^(−sigma²) ≈ 250 (this is equation (99) of Appendix B), while the median (fifty-fifty value splitting the lognormal density in two parts with equal undergoing areas) falls at about N_median = e^mu ≈ 1740. These seem to be smaller values than N = 3500 provided by the classical Drake equations, but it's a wrong impression due to a poor "intuitive" understanding of what statistics is! In fact, neither the mode nor the median are the "really important" values: the really important value for N is the MEAN VALUE! Now if you look at the thin curve in Figure 2 below (i.e. the lognormal distribution arising from the Central Limit Theorem), you see that this curve has a LONG TAIL ON THE RIGHT! In other words, it does NOT immediately go down to nearly zero beyond the peak of the mode. Thus, when you actually compute the mean value, you should not be too surprised to find out that it equals
mean value of N = e^mu · e^(sigma²/2) ≈ 4589.559 ≈ 4590
communicating civilizations now in the galaxy. This is the important number, and it is HIGHER than the 3500 provided by the classical Drake equation. Thus, in conclusion, THE STATISTICAL EXTENSION of the classical Drake equation INCREASES OUR HOPES to find an extraterrestrial civilization!
Figure 2. Comparing the Two Probability Density Functions of the Random Variable N Found (1) Without Resorting to the CLT at All (thick curve) and (2) Using the CLT and the Relevant Lognormal Approximation (thin curve).
Even more so our hopes are increased when we go on to consider the standard deviation associated with the mean value 4590. In fact, the standard deviation is given by equation (97) of Appendix B. This yields
sigma_N = e^mu · e^(sigma²/2) · sqrt(e^(sigma²) − 1) ≈ 11195
and so the expected number of N may actually be even much higher than the 4590 provided by the mean value alone! The "upper limit of the one-sigma confidence interval" (as statisticians call it), i.e. the sum 4590 + 11195 = 15,785, yields a higher number still! (Note: the "lower limit" of the one-sigma confidence interval is ZERO because the lognormal distribution is POSITIVE (or, more correctly, non-negative)). Finally, the reader should note that the thick curve depicted in Figure 2 is just the NUMERICAL solution of the statistical Drake equation for a FINITE number of 7 input factors. Figure 2 actually shows that this curve "is well interpolated" by the lognormal distribution (thin curve), i.e., by the neat analytical expression provided by the Central Limit Theorem for an INFINITE number of factors in the Drake equation. That is, in conclusion, Figure 2 visually shows that taking 7 factors or an infinity of factors "is almost the same thing" already for a value as small as 7.
8. Finding the Probability Distribution of the ET-Distance By Virtue of the Statistical Drake Equation
Having solved the statistical Drake equation by finding the lognormal distribution, we are now in a position to solve the ET-DISTANCE problem by resorting to statistics again, rather than just to the purely deterministic Distance Law (5), as we did in Section 2. This is "scientifically more serious" than just the purely deterministic Distance Law (5) inasmuch as the new statistical Distance Law will yield a PROBABILITY DENSITY for the Distance, with the relevant mean value and standard deviation. In other words, the Distance Law (5) itself becomes a random variable whose probability distribution, mean value and standard deviation must be computed by "replacing" into (5) the fact that N is now known to follow the lognormal distribution.
The important new result is the PROBABILITY DENSITY FOR THE DISTANCE, equation (9) of the present document, which is equation (114) of Appendix B, holding for r ≥ 0.
Starting from this equation, the MEAN VALUE OF THE random variable ET_DISTANCE is computed as
mean value of ET_Distance = C · e^(−mu/3) · e^(sigma²/18) (10)
which is equation (119) of Appendix B, and finally the ET_DISTANCE STANDARD DEVIATION
sigma_ET_Distance = C · e^(−mu/3) · e^(sigma²/18) · sqrt(e^(sigma²/9) − 1) (11)
which is equation (123) of Appendix B. Of course, all other descriptive statistical quantities, such as moments, cumulants etc. can be computed upon starting from the probability density (9), and the result is Table 2 hereafter, that is Table 3 of Appendix B.
Finally, to complete this section, as well as this "introduction to the statistical Drake equation," the numerical values that equations (10) and (11) yield for the input table are determined. They are, respectively:
mean value = C · e^(−mu/3) · e^(sigma²/18) ≈ 2,670 light years (12)
which is equation (153) of Appendix B, and
sigma_ET_Distance ≈ 1,309 light years (13)
which is equation (154) of Appendix B.
Table 3. Summary of the Properties of the Probability Distribution That Applies to the Random Variable ET_Distance Yielding the (average) Distance Between Any Two Neighboring Communicating Civilizations in the Galaxy, assuming they are UNIFORMLY distributed throughout the whole galaxy volume. The distribution is unnamed; the table gives its density, the numerical constant C = cube root of (6 · R_Galaxy² · h_Galaxy) ≈ 28,845 light years related to the Milky Way size, the mean value, variance, standard deviation, all k-th moments, mode, peak value, median C · e^(−mu/3), skewness and kurtosis, and expresses mu and sigma² in terms of the lower (ai) and upper (bi) limits of the Drake uniform input random variables Di.
It is clarifying to draw the graph of the ET_Distance probability density (9):
Figure 3. The Probability of Finding the Nearest Extraterrestrial Civilization at the distance r From Earth (in light years) if the Values Assumed in the Drake Equation are Those Shown in the Input Table. The relevant probability density function is given by equation (9). Its mode (peak abscissa) equals 1933 light years, but its mean value is higher since the curve has a long tail on the right: the mean value equals in fact 2670 light years. Finally, the standard deviation equals 1309 light years: THIS IS GOOD NEWS FOR SETI, inasmuch as the nearest ET galaxy civilization might lie at just 1 sigma = 2670 − 1309 = 1361 light years from us.
From Figure 3 we see that the probability of finding extraterrestrials is practically zero up to a distance of about 500 light years from Earth. Then it starts increasing with the increasing distance from Earth, and reaches its maximum at
r_mode = r_peak = C · e^(−mu/3) · e^(−sigma²/9) ≈ 1,933 light years. (14)
This is the MOST LIKELY VALUE of the distance at which we can expect to find the nearest extraterrestrial civilization.
It is not the mean value of the probability distribution (9). In fact, the probability density (9) has an infinite tail on the right, as clearly shown in Figure 3, and hence its mean value must be higher than its peak value. As given by (10) and (12), its mean value is about 2670 light years. This is the MEAN (value of the) DISTANCE at which we can expect to find extraterrestrials.
After having found the above two distances (1933 and 2670 light years, respectively), the next natural question that arises is: "what is the range, back and forth around the mean value of the distance, within which we can expect to find extraterrestrials with the highest hopes?" The answer to this question is given by the notion of standard deviation that we already found to be given by (11) and (13), sigma_ET_Distance ≈ 1309 light years.
More precisely, this is the so-called 1-sigma (distance) level. Probability theory then shows that the nearest extraterrestrial civilization is expected to be located within this range, i.e. within the two distances of (2670 − 1309) = 1361 light years and (2670 + 1309) = 3979 light years, with probability given by the integral of the ET_Distance density taken in between these two lower and upper limits, that is:
integral from 1361 light years to 3979 light years of f_ET_Distance(r) dr ≈ 0.75 = 75% (15)
In plain words: with 75 percent probability, the nearest extraterrestrial civilization is located in between the distances of 1361 and 3979 light years from us, having assumed the input values to the Drake Equation given by the input table. If we change those input values, then all the numbers change again, of course.
9. The "Data Enrichment Principle" as the Best CLT Consequence Upon the Statistical Drake Equation (Any Number of Factors Allowed)
As a fitting climax to all the statistical equations developed so far, let us now state our "DATA ENRICHMENT PRINCIPLE." It simply states that "The Higher the Number of Factors in the Statistical Drake equation, The Better."
Put in this simple way, it simply looks like a new way of saying that the CLT lets the random variable Y approach the normal distribution when the number of terms in the sum approaches infinity. And this is the case, indeed.
10. Conclusions
We have sought to extend the classical Drake equation to let it encompass Statistics and Probability.
This approach appears to pave the way to future, more profound investigations intended not only to associate "error bars" to each factor in the Drake equation, but especially to increase the number of factors themselves. In fact, this seems to be the only way to incorporate into the Drake equation more and more new scientific information as soon as it becomes available. In the long run, the Statistical Drake equation might just become a huge computer code, growing in size and especially in the depth of the scientific information it contains. It would thus be Humanity's first "Encyclopaedia Galactica."
Unfortunately, to extend the Drake equation to Statistics, it was necessary to use a mathematical apparatus that is more sophisticated than just the simple product of seven numbers.
Appendices
Appendix A: Proof of Shannon's 1948 Theorem Stating That the Uniform Distribution is the "Most Uncertain" One Over a Finite Range of Values. The appendix reproduces the two theorems Shannon proves on pages 36 and 37 of "A Mathematical Theory of Communication" (1948) — that the one-dimensional distribution of fixed standard deviation with maximum entropy is the Gaussian, and that the maximum-entropy probability distribution over any finite interval ai ≤ x ≤ bi is the uniform distribution. (Sections omitted for length; the complete text is at the source.)
Appendix B: Original Text of the Author's Paper #IAC-08-A4.1.4 Entitled "The Statistical Drake Equation," presented at the 59th International Astronautical Congress, Glasgow, 1 October 2008, together with the references. It carries the full derivations behind every equation cited above, including the proof that a finite product of uniform random variables has no analytic form, the lognormal properties table, the distance density function — dubbed the "Maccone distribution" by Paul Davies — and the numerical MathCad example. (Sections omitted for length; the complete text is at the source.)
The way in
https://documents2.theblackvault.com/documents/dia/AAWSAP-DIRDs/DIRD_25-DIRD_An_Introduction_to_the_Statistical_Drake_Equation.pdfDefense Intelligence Reference Document, Acquisition Threat Support, 11 March 2010 (IOD: 1 December 2009), one in the series of advanced technology reports produced in FY 2009 under the DIA Advanced Aerospace Weapon System Applications (AAWSA) Program. Released under FOIA and published by The Black Vault. The preparing office is withheld under 10 USC 424 and the author’s name under FOIA exemption (b)(5)/(b)(6); internal evidence is unusually direct — Appendix B reproduces the author’s own congress paper IAC-08-A4.1.4, printed there under the name Claudio Maccone of the International Academy of Astronautics SETI Permanent Study Group — so the redacted author is inferred to be Maccone, but the released copy itself names no one. Sections 1 through 10, the body of the report, are given here in full. Appendix A (a reproduction of Shannon’s 1948 maximum-entropy proof) and Appendix B (the complete original congress paper, 27 pages in two-column facsimile) are omitted for length; the complete text is at the source. Equations that the released scan rendered unreadable are transcribed in plain notation or noted in place, and the two summary tables of lognormal properties did not survive the scan.
How to cite it
DIA / AAWSAP contractor (2010) DIRD An Introduction to the Statistical Drake Equation. https://documents2.theblackvault.com/documents/dia/AAWSAP-DIRDs/DIRD_25-DIRD_An_Introduction_to_the_Statistical_Drake_Equation.pdf
Where it sits in the curriculum